Choose GLM 5.3 Flash for text, image, and video input; understand required reasoning, structured output, and the per-call price floor.
Where: A supported chat or text-generation model picker → GLM 5.3 Flash.
Key ideas
- A separate model from GLM 5.3: The saved model ID is glm-5.3-flash; the OpenRouter transport uses z-ai/glm-5.3-flash. It is a language-model/orchestrator choice, not an image generator or video renderer. The similarly named GLM 5.3 remains a separate text-only choice.
- Native visual input: Flowmo declares native text, image, and video input for Flash. A video attachment can therefore reach it as video rather than a guessed description. PDF is not declared native: document text is prepared before the request. Audio uses the existing transcription fallback rather than a native-listening claim.
- Reasoning cannot be disabled: The supported picker choices are Auto, Low, Medium, High, and Max. Off is intentionally absent. Max maps to the provider effort supported by this route; it does not make the response length unlimited.
Steps
- Open the model picker on the chat or text-generation surface that will do the work, and select GLM 5.3 Flash rather than GLM 5.3.
- Attach the images or video the answer must actually inspect. For a PDF or Markdown brief, include the document itself and verify the assistant has its text; a private URL alone is not its contents.
- Keep Auto for provider-default reasoning or select a supported effort. Do not send an Off value through an old preset or manually authored MCP payload.
- Read the current “from” price before starting, run a small representative request, and check the answer against the attached evidence.
- For structured automation, request the required schema through the supported tool interface, then validate the returned fields before forwarding them to another node.
Price floor versus final charge
The catalog assigns glm_5_3_flash a base floor of 1 Spark per round. A multi-step agent may make several rounds. Long uncached input and measured output above the included allowances can increase the charge, and reasoning contributes to output usage. The picker’s “from” label is not a fixed price for an entire task.
Structured output and attachments
- Flash supports the bridge’s strict JSON-schema routing and tool calls; do not inherit the older GLM 5.3 structured-output limitation merely because the names share a prefix.
- Image and video attachment capability is read from the model catalog. Text-only models must not receive native visual blocks just because Flash can.
- A text extraction or transcription fallback is additional preparation; it is not equivalent to the model reading the original layout or listening natively.
Current language models and measured pricing
The shared catalog adds Gemini 3.8 Flash, Qwen3.8 Flash, GPT-6 Astra, Claude Haiku 4.5, Claude Sonnet 5, and Claude Opus 5. Saved Gemini 3.5/3.6/3.7 Flash choices now resolve to 3.8. Historical labels remain unchanged, and the visible base round may increase with effort, measured usage, or Astra’s large-context tier.
- Qwen image/video capability is catalogued, but live vision was not verified during the refresh because the upstream route was unavailable.
- Astra/Claude default to OpenRouter for reliable prefix caching; optional Kie-first routing is a server setting.
- The current action quote is authoritative for a paid run.
Troubleshooting
An imported setting tries to disable thinking.
Choose Auto or a supported effort in the current picker. Flash requires reasoning and rejects a provider request that disables it.
The charged amount exceeds 1 Spark.
Check the number of model rounds and input/output usage. The catalog number is the base round floor, not a maximum for long generations.