generate_openai.py, generate_gemini.py, generate_imagen.py, and generate_fal.py each own their provider API details and expose the same basic command shape.
Image-gen design guide
How the image-gen skill works.
The skill is a small comparison-first image generation system: provider adapters create candidates, a local gallery captures preference, and protocol checks keep the pieces interchangeable.
The skill is built from plain scripts with a tested boundary.
There is no provider framework hidden behind the skill. The implementation uses directly executable UV scripts, one shared helper for local mechanics, a review gallery for ranking, and a test suite that checks the adapter interface before paid provider calls happen.
scripts/common.py handles environment loading, output paths, MIME extensions, image dimensions, and generation history. It is the only shared local import for provider scripts.
run_examples.py can produce reviewable comparison manifests. review_gallery.py opens those runs, saves rankings, copies the winner, and supports regeneration.
evaluate_review_gallery.py captures screenshots and checks the gallery behavior, including the image-only default state and non-overlapping metadata rule.
There are two working paths: direct generation and comparison review.
Direct generation is still the simplest path when the desired provider is known. Comparison review is the path for model comparison, ranking, feedback capture, and future training data.
# Direct provider generation skills/image-gen/scripts/generate_openai.py \ --prompt "a quiet app icon for a journaling tool" \ --output /tmp/icon.png # Fixed comparison example, paid only with --generate skills/image-gen/scripts/run_examples.py --generate --example typography-poster # Open the generated run for ranking and feedback skills/image-gen/scripts/review_gallery.py --run-dir skills/image-gen/outputs/examples/<timestamp>
The protocol is the stable surface between providers and tools.
Provider internals can differ. The outer behavior cannot. Runners and review tools depend on a small adapter contract instead of reaching into provider code.
--prompt, --output, --output-dir, --model, --check, and --protocol.--protocol prints machine-readable metadata without requiring API keys, provider credentials, or generation spend.data/generations.jsonl unless IMAGE_GEN_DISABLE_HISTORY=1 is set.The tests check the interface, not only the implementation.
The default test suite stays free. It compiles scripts, checks executable UV metadata, validates the provider protocol surface, exercises --protocol JSON with no API keys, verifies history logging, and checks that generated example manifests are review-gallery compatible. The focused check_protocol.py wrapper runs the adapter-contract checks directly and is part of PR validation for image-gen changes.
A provider can be internally different and still be externally interchangeable.
No-key protocol command. check_protocol.py calls provider scripts with --protocol using the local Python interpreter, so provider SDKs and credentials are not needed.
Static guardrails. Tests still scan options, top-level functions, local imports, history calls, and final stdout behavior.
Manifest handoff. The example runner now records candidate entries in the shape that the review gallery already consumes.
Paid calls separate. Real generation remains behind explicit --generate flags and provider key checks.
The skill accumulates local evidence as it is used.
Generated images, manifests, evaluation screenshots, rankings, regeneration events, and provenance logs are intentionally local and gitignored. The tracked repo carries the tools and reusable assets; the ignored directories carry private usage history. This is a documented local convention for this skill rather than a general Agent Skills standard.
outputs/
Generated images, example comparison runs, review manifests, winner copies, test runs, and visual evaluation screenshots.
data/
Local JSONL history including generations.jsonl, rankings.jsonl, and regenerations.jsonl.
assets/
Tracked reusable inputs such as the gallery template, prompt packs, masks, and safe reference materials.
references/
Tracked maintenance notes, model matrices, this design guide, and the storage policy explaining why mutable data is separated from reusable skill resources.
The remaining work is to make comparison the normal prompt path.
The foundation is now in place: provider scripts advertise their protocol, tests enforce the interface, and example manifests can feed the review gallery. The next step is a single prompt-first command that runs enabled providers, opens review, waits for save, and prints only the winning image path.
compare_prompt.py
Generate one candidate per enabled provider for arbitrary prompts, write a comparison manifest, launch the gallery, and return the selected winner.
Run history
Add comparison_runs.jsonl so the skill can analyze prompts, model choices, rankings, comments, and winners over time.
Standalone export
Add a stricter checker or export path for fully copy-pastable provider files after the protocol has stabilized.
Small adapters
Keep provider differences visible inside provider files.
Protocol first
Make cross-script behavior explicit and testable.
Image-first review
Do not let controls, metadata, or commentary dominate the images being judged.