Agent workflow
Model discovery, references, durable requests, and saved files.
Read as Markdown ↗Build integrations
Read llms.txt to find focused pages, or llms-full.txt for the entire reference. Download OpenAPI for schemas. Each page has a Markdown URL, for example /markdown/api/image.md.
Use server-side bearer authentication. Keep secrets out of client code. Discover model IDs through GET /api/generate/models and constrain inputs to that model's capabilities. Model discovery is not a quote or full per-model schema endpoint.
A useful durable client record contains the originating API-key identifier (not its secret), idempotency key, exact request body, submission timestamp, returned task/request ID, and current state. Persist this before dispatch. Retain it across reloads and process restarts.
Generate files
Before spending, establish the user's requested output and numeric Gem ceiling. Submit only the approved quantity and settings. Include maxGems and never silently increase it. Report PRICE_CHANGED so the user can reduce settings or revise their budget.
Upload owned files only when needed. Image/video/audio references use asset IDs; private cloning uses the separate voice upload and the returned voice profile ID. Confirm voice rights with the user before setting rightsConfirmed=true.
For media, poll with backoff and download successful outputs. Read the task again for expired URLs. For text or extraction, replay the exact request identity to obtain retained output. In uncertain cases, keep the identity and stop duplicate dispatch.
Verify without paid inference
For integration development, start with mocked API responses, schema checks, and authenticated read-only model/voice discovery. Do not treat successful build/typecheck as proof of a real provider generation. Only perform a live generation test when the user has approved that spending and a numeric budget.
After a live media call, report task ID, terminal state, quoted/charged/released Gems, successful files, and remaining validation limits. Text/extraction responses do not include the media task's Gem breakdown; do not invent charges from token counts.