Agentic workflows and custom skills

Built at Formative Minds

How I orchestrated multi-agent pipelines to facilitate software and media production with Claude Code.

role
Designed workflows, custom skills, and the review and spend gates.
stack
Claude Code · MCP · Higgsfield · ElevenLabs · FFmpeg · Python

Tasks are routed to the model suited to the role (Fable orchestrates and reviews, Opus executes), with tests and human review before anything irreversible.

Failures are written down and turned into standing rules, so the same mistake doesn't come back.

The agentic development workflow A custom skill sets out the steps and gates for a job. Fable orchestrates and reviews; Opus subagents carry out the plan, tests first. Nothing irreversible happens without human review. Each checkpoint records what went wrong, and the lesson becomes a standing instruction that every later session reads.
The generative AI pipeline Every image, video or audio generation starts with a free cost check and an explicit approval. Higgsfield is driven through MCP tool calls, ElevenLabs through its REST API. A rejected result is diagnosed before anything is re-generated, and the lesson goes into the playbook. Locked results are logged, reconciled against the real balance, and assembled locally for free.

16 custom skills

A skill is a written procedure an agent follows when a job comes up: when to use it, the steps, and the checks before anything is final.

Half of these skills serve media production. The game is set in its own story world, so characters, art and films have to stay consistent. Every creative output is reviewed by a person before it's used.

Engineering discipline

  • checkpoint Saves progress: devlog entry, commit, push.
  • green Runs every test suite and reports whether the build is green.
  • audit Runs a full code-review loop and records the findings.
  • docs-tidy Reconciles the docs with what the code actually does.

Product and deployment

  • kiosk-ship Builds and deploys the kiosk app.
  • box-kit Assembles, tests and ships a new kiosk box.
  • season Creates, switches or audits the game's active season.
  • deck-lab Checks whether target cards can actually be built from the deck.

Generative AI production

  • gen Guards every image and video generation: preflight, approval, logging.
  • spend Reconciles credit spend against the production ledgers.
  • art Generates and refines game art to the style guide.
  • idle-avatar Produces looping idle animations for character avatars.
  • cutout Cuts characters out of their backgrounds and repairs bad cutouts.
  • monster-intro Plans and produces a monster intro video, shot by shot.

Worldbuilding

  • lore Creates and validates entries in the story world's database.
  • blight Adds or edits a monster entry.

deep dive into 3 important custom skills

checkpoint

The context-engineering loop: every save leaves a written record the next session reads.

trigger
"checkpoint", "commit and push", "sync everything up"
steps
  1. Run the test gate if code changed
  2. Write a devlog entry: what shipped, why, decisions, what's still owed
  3. One atomic commit and push
  4. Report anything skipped, and why
gates
  • Tests green before committing
  • A failed step is reported, never hidden
output
A devlog entry and a pushed commit that point at each other

spend

Cost governance: generation credits are checked against what the ledgers claim.

trigger
Checking credits, reconciling spend, or finding unlogged generations
steps
  1. Read the real balance and transaction history from the API
  2. Collect job ids from the production ledgers
  3. Diff both ways: spend with no ledger row, rows with no spend
  4. File the missing rows in each ledger's format
gates
  • The API is ground truth, not the docs
  • An ambiguous match is asked about, never guessed
output
Ledgers whose arithmetic closes against the real balance

deck-lab

Where the maths meets the agents: can every target card really be built?

trigger
"check solvability", "can this level build 42?", "what if we add a target?"
steps
  1. Enumerate every possible hand for a level
  2. Count how many hands can build each target
  3. Score candidate cards against the rule
gates
  • Exhaustive, never sampled
  • A card joins a level's pile only if at least 95% of hands can build it
output
An exact solvability table per level, and a ruled card list

build notes

Built with Claude Code multi-agent workflows: I designed the architecture, owned the decisions, and verified the output; agents wrote most of the code.