Summary
AI period brief · Last 30 days
Docs RAG citation quality slipped after the Jul 30 traffic spike, citation F1 fell to 0.79 against the 0.82 gate, concentrated in multi-hop questions.
docs.rag · 180 samples
Reasoning-model spend is up 38% week over week, driven by code.review.diff volume with no cheap path for small diffs.
+$184/day run-rate
support.refund.policy on assist-fast holds a 91 score at $0.003/run: the strongest cost/quality position on the frontier this period.
frontier top decile
Runs (sample)8.2%
320
278 ok
Spend (sample)12%
$1.6K
all models
Avg eval score1.4
76
scored runs only
Avg latency40ms
3.2s
all statuses
Error rate0.4pp
13.1%
error+timeout+filtered
Pareto cellsfrontier
5
of 72 scored
Spend vs run volume
Northline AI ops · daily
Spend $Runs
Jul 12 assist-large default for supportJul 21 Eval suite v4 (groundedness)Jul 30 Traffic spike · promo weekAug 2 Budget alert · reasoning model
Runs by status
320 traces
278
ok
21
error
14
timeout
7
filtered
Slowest ok runs
by latency
| Prompt | Model | Lat |
|---|---|---|
| docs.triage.ab | assist-reasoning | 9.1s |
| docs.discovery.shadow | assist-reasoning | 9.0s |
| sales.battlecard.lite | assist-reasoning | 9.0s |
| docs.triage.ab | assist-reasoning | 8.9s |
| code.battlecard.v2 | assist-reasoning | 8.7s |
| sales.battlecard.lite | assist-reasoning | 8.4s |
| support.tests.exp | assist-reasoning | 8.3s |
Top prompts by volume
in sample window
| Prompt | Runs | Spend | Score |
|---|---|---|---|
| code.battlecard.v2 | 10 | $55.03 | 75 |
| support.refund.lite | 8 | $23.38 | 65 |
| support.discovery.shadow | 8 | $2.49 | 63 |
| sales.faq.canary | 7 | $9.41 | 66 |
| docs.discovery.v2 | 7 | $44.25 | 87 |
| docs.triage.ab | 7 | $76.76 | 89 |
| docs.discovery.shadow | 7 | $50.25 | 90 |
| sales.rag.v1 | 7 | $2.83 | 59 |
Sample budget pace
$42K / month
55%
pace
Sample spend$1.6K
Avg score76
Error rate13.1%
Month budget$42K
Pace is illustrative from the demo sample (320 runs), not live billing.
Root-caused alerts
symptom → cause → fix
Do this next
scored from frontier · quality lift ÷ cost
#192
Move Docs hard queries → assist-large
Frontier: +6 score for +$0.09/run on 15% traffic
$890.00 linked spend
#288
Cheap path for small code diffs
Cut reasoning spend ~30% without recall drop on low-sev
$184.00 linked spend
#380
Fix retrieve timeout on triage
Error cluster · 12 runs · no quality cost but SLA
$48.00 linked spend
#474
Pin sales voice glossary
Recover email eval gate with one prompt commit
$40.00 linked spend