WeGenDocs
Generation API

Text and streaming

OpenAI-shaped chat completions with durable replay.

Read as Markdown ↗

POST /api/generate/text/v1/chat/completions

Requires generate scope and Idempotency-Key. Discover a reviewed text model before submitting. This is a subset of the OpenAI chat-completions shape; it is not a Responses, tools, or multimodal chat endpoint.

FieldTypeLimits / default
modelstring, requiredEnabled reviewed text model ID, 1–150 characters.
messagesarray, required1–60 objects with role and string content.
messages[].rolestringsystem, user, or assistant.
messages[].contentstringUp to 30,000 characters each.
streambooleanfalse.
max_tokensinteger1–8,192; default 4,096.
temperaturenumberOptional, 0–2.
top_pnumberOptional, 0–1.
response_formatobjectOptional {type: "text"} or {type: "json_object"}.
maxGemsintegerOptional conservative admission ceiling, 0–1,000,000.

Request bytes are bounded to 160,000; serialized conversation bytes to 128,000. Unknown fields, tools, provider-routing overrides, remote references, and array-valued message content are rejected.

Non-streaming example

curl https://wegen.art/api/generate/text/v1/chat/completions \
  -H "Authorization: Bearer $LETSGEN_API_KEY" \
  -H 'Content-Type: application/json' \
  -H 'Idempotency-Key: my-project-text-001' \
  --data '{
    "model": "TEXT_MODEL_ID_FROM_DISCOVERY",
    "messages": [{"role": "user", "content": "Write a short scene about a moonlit harbor."}],
    "max_tokens": 256,
    "maxGems": 10
  }'

HTTP 200 returns id, object: "chat.completion", Unix created, model, choices with assistant content and finish_reason, and usage token counts.

Streaming

Set stream: true and consume Content-Type: text/event-stream. Use curl -N to disable buffering. SSE data: frames contain chat.completion.chunk objects. Accumulate choices[].delta.content; the final successful chunk includes usage, followed by data: [DONE].

A stream can also contain {error: {code: "REQUEST_UNCERTAIN", message}}. Treat an error or a stream ending without [DONE] as incomplete; never assume HTTP 200 alone proves success. Disconnecting does not cancel provider spending.

Budgets and replay

The API reserves a conservative whole-Gem maximum from the reviewed rates, input size, and requested output limit before inference. Measured usage settles once through the account LLM Gem allowance and consent-gated paid Gems. The charge cannot exceed the admitted reservation.

Repeat the exact body, including stream, with the same originating key and identity to retrieve a saved response. A completed streaming replay may return the whole saved answer in one chunk instead of reproducing the original chunk timings. Pending or ambiguous responses return 409 REQUEST_UNCERTAIN; they never dispatch again. See Budgets and retries.

On this page