Metrics Dashboard
The Metrics dashboard provides execution analytics, helping you understand how efficiently Trinity is building your project.
Accessing Metrics
Click Metrics under the Workspace section in the sidebar. It sits alongside Activity and Recaps — the pages that read the whole workspace and need no project in focus. Data updates in real time as execution events fire on the server. If the WebSocket push channel is unavailable, Trinity checks for changes every 30 seconds until it is back.
Filter Bar
The page has a two-row filter bar at the top.
Row 1 — scope cascade: Project → Release → PRD.
- Project — defaults to All projects, which aggregates every project in the workspace, plus the spend that belongs to no project at all (the model-setup chat and other assistant conversations, which are yours rather than any project's). Pick one to narrow the charts to it. This is a filter on this page only: it never changes the app's active project, so the sidebar, Dashboard, and Stories keep looking at whatever they were.
- Release — defaults to All releases. Pick a specific release to scope every story-derived chart to it. Available once you have filtered to a specific project — Release has no meaning under "All projects" — regardless of which project is active elsewhere in the app.
- PRD — defaults to All PRDs. If a release is selected, the PRD list narrows to PRDs in that release; with "All releases" it lists every PRD in the project. The PRD filter scopes the Detail tab, under the same condition as Release.
Row 2 — period: the date-range row, described below.
Two tabs ignore the release and PRD selection and cover whatever the project filter is showing: Workers (workers process jobs from any release) and AI Usage (planning and provisioning calls happen before any one release exists, so scoping them would misrepresent where the budget went).
Period Filter
A Period row beneath the scope cascade lets you scope every chart by date. The page opens with 30d active, so the first thing you see is the last 30 days of activity rather than the project's entire history:
- All time — no date filter; still available as an explicit choice
- 1d / 7d / 30d / 90d — rolling windows ending on your current day, in your configured timezone
- Custom from / to date pickers — pick any range; switches the active preset to "custom"
Dashboard Tabs
Overview
The overview tab shows North Star metrics — the most important indicators:
- Success Rate — percentage of stories that complete successfully (vs. failing)
- First-Pass Rate — percentage of stories that pass on the first attempt without needing retries
- Cost per Merged Story — average token cost to complete and merge a story
- Average Cycle Time — how long stories take from start to completion
These metrics help you evaluate the overall health of your execution pipeline.
Pipeline
Visualizes the story execution funnel:
- Pending → Claimed → Running → Complete/Failed — how stories flow through the system
- Queue Wait Time — how long stories wait before a worker picks them up
- Gate Wait Time — how long stories spend waiting at execution gates
- Retry Distribution — how many attempts stories needed, as a histogram
- Failure Reasons — breakdown of why stories fail
- Agent Performance — per-phase (analyst, implementer, auditor, documenter) totals, accepted and rejected counts, success rate, and average duration. A low auditor success rate, for instance, suggests the implementer is producing code that needs heavy revision.
Use this tab to identify bottlenecks. Long queue wait times suggest you need more workers. Long gate wait times mean you should check for pending approvals more frequently.
Cost
Token usage and cost tracking:
- Total Cost — cumulative spend in USD
- Cache Hit Rate — how much of your input came from cache, with the underlying counts
- Output Tokens — total tokens generated
- Wasted — spend that went to story runs that failed
- Token Breakdown — input, output, cache-read, and cache-write tokens side by side
- Daily Cost Trend — how much you're spending per day
- Cost by PRD — stories, completion, tokens, and average cycle time broken down per PRD
Every AI operation Trinity runs is recorded, and wasted tokens from failed story runs are counted separately so you can see efficiency.
Unpriced calls. Trinity prices each call from its own rate card for the models it ran on, using the rates that were in force when the call happened. A call on a model with no rate card counts as unpriced: its spend is missing from the total rather than estimated in it. Whenever there are any, the Total Cost card flags them instead of showing its usual description, so a headline figure never quietly reads as exact.
Because prices are worked out when you look rather than stored when the call ran, correcting a rate card updates every past figure on this page — there is nothing to re-import or rebuild. Recaps are the deliberate exception: a recap is a written record of a day, so it keeps the figure it was built with.
Users
A per-member breakdown of spend and token usage for everyone who has run AI in the project, sorted by cost:
- Per-Member Table — each member with their cost, input / output / cache tokens, and call count
In a workspace of one this is just your own runs.
Workers
Worker pool health:
- Utilization — percentage of time workers are busy vs. idle
- Total / Busy / Idle — current worker status breakdown
- Retry Distribution — how often stories need to be retried
- Stale Jobs — jobs that have been running too long
- Per-Worker Stats — individual worker performance
AI Usage
Volume, cost, and reliability by model and operation:
- Headline stats — Total AI Calls, Total Tokens (input and output), Total Cost, Success Rate, and Total Duration
- Usage by Category — calls, successes, tokens, and duration grouped by category (planning, execution, and so on)
- Usage by Operation — a per-operation breakdown you can filter by model and category: calls, success rate, cost, total / input / output / reasoning tokens, and average duration. Reasoning tokens are what a model spent thinking before it answered — reported per turn by the Codex engine, and blank for engines that don't report it.
- Daily AI Usage Trend — calls, tokens, and errors per day
When any calls could not be priced, a banner above the tables names how many and which models were affected, and every cost figure here carries the same caveat as the Cost tab's headline.
Per-phase agent success and duration (analyst, implementer, auditor, documenter) live on the Pipeline tab's Agent Performance table, not here.
Releases
Per-release cost breakdown showing the full picture for each release:
- Summary Cards — total releases, total cost, average cost per release, total duration
- Per-Release Table — each release with name (and its order number), status, PRD count, story count, story cost (all stories in the release's PRDs), release cost (SEO audit, preflight, etc.), total cost, tokens, and duration
Story cost aggregates all AI events from stories belonging to the release's PRDs. Release cost covers the release's own execution phases. The total is both combined.
Detail
Per-PRD rollups, filtered by the PRD picker in the top filter bar:
- Completion Percentage — how far along each PRD is
- Token Usage — cost per PRD
- Cycle Time — average story duration per PRD
- Story Breakdown — status distribution for each PRD
Understanding the Metrics
Success Rate
A healthy project typically has a success rate above 80%. Lower rates suggest:
- Stories are too vague (improve acceptance criteria)
- Dependencies are missing (stories fail because required code doesn't exist)
- External services aren't configured (missing secrets)
First-Pass Rate
This measures how often stories succeed without retries. A high first-pass rate (>70%) indicates:
- Good planning — stories are well-scoped
- Clean codebase — agents can work effectively
- Accurate difficulty ratings — appropriate resources are allocated
Cost per Story
This varies significantly by difficulty:
- Lower-difficulty stories: lower cost — they tend to run on the standard tier
- Higher-difficulty stories: higher cost — the Calibrator tends to route them to the reasoning tier
- Checkpoints: highest cost (multi-pass audits)
Difficulty itself is a read-only descriptor and doesn't directly set the model tier — the Calibrator judges each story's work and picks the tier, so these are tendencies, not a fixed formula.
Cycle Time
Average time from story start to completion. Affected by:
- Story difficulty and surface area
- Number of reviewers on the story's audit
- Gate wait time
- Worker availability
Timezone Handling
The database stores all timestamps in UTC. The metrics dashboard converts to your local timezone for daily groupings, so the "Daily Cost Trend" chart reflects your actual days.
Tips
- Monitor the pipeline tab after starting execution — it shows real-time flow
- Check cost trends weekly — catch unexpected spending spikes early
- Use worker health to tune parallelism — if utilization is consistently low, reduce workers; if queue wait is high, add more
- Compare PRD rollups — later PRDs should ideally show better metrics than earlier ones