AI Model Configuration

Trinity uses a tiered model system to balance quality and cost across different operations. You can configure which models are used at each tier.

Model Tiers

Trinity has four tiers, each suited to different types of work. A tier's picker always offers its own rung and every rung above it — you can assign a smarter model than a tier strictly needs, never a dumber one. A provider with nothing at a tier's level falls back to its nearest rung instead of disappearing from the picker: its best model when it has nothing that high, its weakest when it has nothing that low.

Frontier Tier

The most capable models available — each vendor's headline release. Settings carries a standing Frontier Model row like every other tier, but it's opt-in only: no operation, calibrator pass, or default cascade routes here automatically — you reach it only by hand — see Opt-in Frontier below.

Minimum intelligence level: 4 (Fable-class). Only Anthropic, OpenAI, and Sakana catalog a model this strong.

Reasoning Tier

The most capable configurable tier, used for tasks requiring deep thinking:

  • Complex or security-critical story implementation
  • Full PRD generation (architect, story writer, dependency mapper, and calibrator phases)
  • Codebase audits
  • Architecture analysis

Minimum intelligence level: 3 (Opus-class and above).

Standard Tier

The everyday workhorse, used for routine tasks and bounded judgment calls alike:

  • Routine story implementation
  • Analyst and implementer phases for everyday stories, and the package-mapper PRD phase
  • Story auditing
  • Onboarding Q&A, documentation generation, roadmap section generation, PRD editing

Minimum intelligence level: 2 (Sonnet-class and above).

Micro Tier

Mechanical, low-intelligence tasks:

  • Classification and scoring
  • Preflight checklists
  • Recap search triage
  • Status checks

Minimum intelligence level: 1 (anything catalogued clears this floor).

Providers

Trinity supports ten AI providers:

Anthropic

The primary provider:

  • Claude Fable 5.1 — frontier-class, 1M-token context; opt-in only (see below)
  • Claude Opus 5.5 — reasoning-class, 1M-token context; default for the reasoning tier
  • Claude Sonnet 5 — standard-class, 1M-token context; default for the standard tier
  • Claude Haiku 4.5 — micro-class, 200k context; default for the micro tier

DeepSeek

Alternative provider, V4 generation:

  • DeepSeek V4 Pro — standard-class, 1M-token context
  • DeepSeek V4.1 Flash — micro-class, 1M-token context

Moonshot (Kimi)

  • Kimi K3 — reasoning-class, 1M-token context
  • Kimi K2.7 Code — standard-class, 262k context; the standard- and micro-tier pick for Moonshot

Z.ai (GLM)

Zhipu GLM models via Z.ai's Claude Code integration:

  • GLM 5.3 — standard-class, flagship agentic coder, 1M context; Z.ai has nothing above standard-class, so this is also its reasoning- and frontier-tier pick
  • GLM 4.6V — standard-class, vision-capable variant, 128k context
  • GLM 5.3 Flash — micro-class, 1M context; the micro-tier pick for Z.ai

Qwen (Alibaba Cloud)

Qwen3 family via Alibaba DashScope's Claude Code integration:

  • Qwen3.8 Max — standard-class, ~1M context, native thinking; Qwen has nothing above standard-class, so this is also its reasoning- and frontier-tier pick
  • Qwen3.7 Plus — standard-class, multimodal agent flagship
  • Qwen3 Coder Next — standard-class, agentic tool calling, 200k context
  • Qwen3.8 Flash — micro-class, ~1M context; the micro-tier pick for Qwen

Xiaomi (MiMo)

MiMo models via Xiaomi's Claude Code integration:

  • MiMo V2.6 Pro — standard-class, flagship reasoning model
  • MiMo V2.6 Flash — micro-class, 1M context; Pro's cheap sibling, and the micro-tier pick for Xiaomi

Ollama

Local model support for private, on-machine execution:

  • Qwen3 Coder Next — standard-class, code-focused
  • Qwen3.6 27B — micro-class, current-gen general/coding model
  • GLM 4.7 Flash — micro-class, lightweight MoE (30B total / ~3B active) general-purpose model
  • Qwen 3.5 9B — micro-class, small general-purpose model

OpenAI (via Codex)

GPT-6 and GPT-5.6 models served through the Codex harness (authenticated via the Codex CLI, not an API key). Codex answers all four tiers with four distinct models:

  • GPT-6 Astra — frontier-class, 922k context; the Codex harness's opt-in frontier pick
  • GPT-6 Sol — reasoning-class, 922k context; the Codex default for the reasoning tier
  • GPT-5.6 Terra — standard-class, 922k context; the Codex default for the standard tier
  • GPT-6 Luna — micro-class, 922k context; the Codex default for the micro tier

xAI (via Codex)

Grok, served through the Codex harness as a custom provider — it needs an xAI API key (unlike OpenAI, which uses the Codex CLI login):

  • Grok 4.7 — reasoning-class; xAI catalogs nothing above or below this level, so this model serves every tier through nearest-rung resolution

Sakana AI (via Codex)

Fugu, served through the Codex harness as a custom provider — it needs a Sakana API key from console.sakana.ai (like xAI, and unlike OpenAI, which uses the Codex CLI login).

  • Fugu Ultra — frontier-class, 1M-token context; opt-in only, Sakana's headline model — and, since Sakana catalogs nothing below this level, also the pick for the reasoning, standard, and micro tiers on a Sakana-configured install
  • Fugu Cyber — frontier-class, 1M-token context; a specialist for deep root-cause work on large codebases. It shares Ultra's intelligence level but tier routing never resolves to it — you reach it only by pinning it on a story's step configuration

Fugu models are not available in the EU/EEA. Reaching Fugu Cyber takes two things beyond the API key above: Sakana must approve an access request — submit your use case and verified contact details at console.sakana.ai and wait for their manual review — and the key's Billing mode must be set to Pay as you go on that same console, since Cyber isn't included in Sakana's subscription tiers. Picking Fugu Cyber without both set up doesn't work.

Configuring Models

  1. Navigate to Settings
  2. Find the AI Models section
  3. For each tier, pick a harness → provider → model cascade:
    • Harness — the CLI/runtime that actually executes the agent. Choose Claude Code (which talks to every Anthropic-compat provider below — Anthropic, DeepSeek, Moonshot, Z.ai, Qwen, Xiaomi, Ollama) or Codex (OpenAI's GPT-6 and GPT-5.6 models, xAI's Grok, and Sakana's Fugu models).
    • Provider — which vendor runs the model.
    • Model — the specific model within that provider.

Switching providers restores the last model you picked for that (harness, provider) pair, so you can A/B between two providers without losing your selection. Settings are stored as harness:provider:model strings (e.g., claude-code:anthropic:claude-opus-5-5).

Picking a harness commits you to having its CLI on the machine. Trinity checks that on the way into a workspace, and again when you open a project whose own tiers name a different one — see Prerequisites → When Checks Happen. The check is a screen with two exits, not a dead end: install the CLI inline, or let Trinity move the affected tiers onto the engine you already have. That second exit writes wherever the setting came from — your workspace settings if the project inherited it, your per-project overrides if the project set it for itself — so a workspace-wide mistake gets fixed workspace-wide rather than once per project.

Defaults

Settings has a configurable row for each of the four tiers:

Tier Default Model
Frontier Claude Fable 5.1
Reasoning Claude Opus 5.5
Standard Claude Sonnet 5
Micro Claude Haiku 4.5

Frontier's row works exactly like the other three — it's just never reached automatically (see Opt-in Frontier below). A manual pick or per-story override that asks for it resolves to whatever you've configured there, or Claude Fable 5.1 if you haven't touched it, or, on a provider with no frontier-class model, that provider's best available model.

The first entry in each row's model list is the inherit entry, and it reads two ways depending on the surface. On a settings tier it names the layer above — Inherit workspace, Inherit default — because the value there is somebody else's to change. On a project's Model Overrides it names the resolved model instead — Inherited (<model>), the model above — because what a project reader wants to know is what will actually run. Either way, choosing it leaves the row unset so the layer above keeps answering.

Retired models appear greyed out in the picker and aren't selectable. Stored settings that reference them keep working (and appear correctly in historical metrics) until you change them.

Only a real model can be saved

Every place Trinity stores a model — your own tiers for a workspace or for a single project, a workspace's defaults, a project's Model Overrides, a story's per-step picks, a release's automation overrides, and a one-conversation chat override — checks that the model is one Trinity actually carries before saving it. A name Trinity doesn't recognise is refused at the moment you save, and the message names both the row that's wrong and the value you tried to save, rather than the value being stored and the run failing on it later. Clearing a tier back to the default is still saved: that's an empty value, not a bad one.

The pickers only ever offer models Trinity carries, so this bites on values that don't come from a picker — a model set through the API, or one an agent proposed for a story's step. The check runs on the server that stores the row, so it holds however the value got there. Retired models are unaffected: they stay in the catalog (just hidden from the pickers), so a setting that names one still saves.

Tier Resolution

Most operations have a fixed tier. Two cases differ:

  • Story execution — the implementation phase uses the model tier set on the story (its tier execution setting). The calibrator authors this at reasoning or standard — it never assigns micro or frontier on its own — but you can override any step's tier by hand on the story's Execution Settings card, including up to frontier. A story flagged for reasoning therefore costs more than a standard one, even though both go through the same pipeline.
  • Planning pipeline — the architect, story writer, dependency mapper, and calibrator phases all run at reasoning; the package mapper phase runs at standard.

Tier Default Resolution

Trinity picks which model to use in this priority order (highest wins):

  1. Explicit — if the caller passes a specific model, that wins
  2. Entity / job override — a model carried on a story's or release's automation overrides, or one passed when starting a run. There's no per-story/release model picker in the UI (the story Automation tab covers PR/merge toggles, not model choice), so this layer is set programmatically and rarely by hand
  3. My project overrides — your own model choice for this project
  4. Project default — the project's configured tier model
  5. My workspace settings — your own model choice for this workspace
  6. Workspace settings — the workspace's configured tier model
  7. TIER_FALLBACK_MODEL_ID — built-in Anthropic defaults used when no layer above has been configured

If the model a layer names is one Trinity no longer offers, or one below that tier, the tier runs its default model from the same provider instead. The model picker shows the stored model as unavailable until you pick another, and saving your other settings still works in the meantime.

This is resolution-time only — it's not a runtime "retry with a fallback on failure." If a model call fails, the operation surfaces the error (the caller handles retries, usually by going through the feedback pipeline or job-level retry).

Per-thread overrides in chat

Inside any AI conversation — Help Chat, the Architect, and the runtime agent on a project's Runtime page — a model button in the composer opens a picker that sets the model — and, for models that support them, the reasoning effort and the answering speed — for that one conversation. The override sticks to the thread until you change or reset it, and applies only there: it doesn't touch your project or workspace defaults, or the models a run uses for its pipeline steps.

The picker always opens on what the conversation will actually run — harness, provider, model, effort and speed are each highlighted whether that came from an override you set or from your configured defaults. There is no blank row waiting to be filled in and no "you haven't chosen yet" state: what's lit is what the next reply uses. Apply pins exactly what you see, so the conversation keeps that selection even if your defaults later change. Reset to default appears only once the conversation carries an override of its own, and removes it so the conversation goes back to inheriting from the resolution order above.

Speed

Some models answer at more than one speed. Where one does, the picker shows a Speed row: Standard — the vendor's normal queue, and what every conversation uses unless you say otherwise — or Fast, which returns the same model's answer sooner and is billed at the vendor's premium rate for that tier. The row appears only for models whose vendor actually sells a faster tier; today that is the GPT models Trinity runs through Codex, where the vendor calls it Fast. Anthropic, xAI and the local providers publish no such tier, so no row appears for them and switching to one of those models drops a Fast setting back to Standard.

One thing worth knowing before you pick it: Fast is a request, not a guarantee. Under heavy load the vendor may answer a Fast request at standard speed and charge standard rates for it, and it doesn't tell Trinity when it does — so a conversation left on Fast can occasionally cost a little less than Trinity's metrics show. A tier the model doesn't actually offer is the other direction and is handled: Trinity records what the vendor confirmed, not what you asked for, so the cost you see is never inflated by a request that was turned down.

The Default tag in the model list marks the model your settings resolve to for this conversation — the one you'd get with no override at all. It moves as your project or workspace choice changes, so the tag and "reset takes me here" are always the same model.

These pickers are scoped to the standard tier, but — per the "own rung and above" rule — that includes every model above it too, frontier-class ones included. So picking Claude Fable 5.1 or GPT-6 Astra for one conversation needs no separate step; you pick it the same way you'd pick any other model. A conversation whose resolved model sits below the tier's bar — a retired model a setting still names, say — is listed anyway, so the picker always shows you what you're running rather than going blank on a model it wouldn't have offered.

Opt-in Frontier

Frontier has a standing default like every other tier (Settings → AI Models → Frontier Model), but it's never assigned automatically — no operation, calibrator pass, or default cascade routes to it on its own. You reach it only by hand, in any of these ways:

  • A manual model pick in Help Chat, the Architect, or a project's Runtime chat.
  • A manual model pick in a story's Execution Settings card, overriding a single step's tier to frontier — this resolves to your configured Frontier model, the same way overriding a step to reasoning resolves to your configured Reasoning model.
  • Setting your standing Reasoning model (Settings → AI Models) to a frontier-class model such as Claude Fable 5.1 — since a tier's picker always allows a smarter model, this makes every reasoning-tier operation run at frontier by default, without touching any per-step or per-turn setting.

This is deliberate: frontier is the most expensive tier, so it never turns on without someone explicitly choosing it.

Effort Levels

Recent Anthropic and OpenAI (Codex) models, plus Grok, accept an effort parameter that controls how hard the model thinks before answering. Each model supports a different subset of the ladder:

  • GPT-6 Astra, GPT-6 Sol, GPT-5.6 Terra (Codex) — 6-level: low | medium | high | xhigh | max | ultra
  • Claude Fable 5.1, Claude Opus 5.5 — 5-level: low | medium | high | xhigh | max
  • GPT-6 Luna (Codex) — 5-level: low | medium | high | xhigh | max — the one OpenAI model without ultra
  • Claude Sonnet 5 — 4-level: low | medium | high | max (no xhigh)
  • Grok 4.7 (Codex) — 4-level: low | medium | high | xhigh
  • Claude Haiku 4.5, and every DeepSeek / Moonshot / Z.ai / Qwen / Xiaomi / Ollama model — no effort parameter; the harness skips injecting it

Trinity's harness clamps every requested effort to what the chosen model supports, so an xhigh request against Sonnet 5 lands at high. Effort requests get recorded in the ai_events table for metrics.

Timeouts

Trinity times AI calls by the kind of work they do, not by which model tier resolves them — a quick classification call and a full PRD generation are timed independently of whether they land on reasoning, standard, or micro. Individual calls range from about 10 minutes (classification, scoring) up to 2 hours (full PRD generation, story implementation, auditing). A queued story or release run also carries its own governor, independent of both: it's killed only after 20 minutes with no filesystem progress in its worktree, or an absolute 4-hour ceiling — whichever comes first — so a legitimately long phase (a slow compile, a big test suite) is never cut off just for taking a while.

Cost Considerations

Model costs vary significantly:

  • Frontier tier is the most expensive, and never turns on by itself — you only pay for it when you've deliberately opted in
  • Reasoning tier is the next most expensive — use it where quality matters most (the calibrator assigns it to complex or security-critical stories, and you can set a story's tier yourself on its Execution Settings card)
  • Standard tier offers the best quality-to-cost ratio for most work
  • Micro tier is very cheap, used for mechanical classification

The Metrics dashboard tracks token usage and cost by operation, helping you understand where your budget goes.

Tips

  • Start with defaults — the default configuration is well-balanced for most projects
  • Use DeepSeek for cost savings — if you have a DeepSeek API key, using it for the standard or micro tier can reduce costs significantly
  • Use Ollama for privacy — local models keep all data on your machine, but expect slower execution and potentially lower quality
  • Monitor the cost tab — check Metrics → Cost to understand your spending patterns before making changes
  • Don't downgrade reasoning — the reasoning tier handles your most complex stories; using a less capable model here leads to more failures and retries, which can cost more in the end