Copilot/Overview

Summary

AI period brief · Last 30 days

  • Docs RAG citation quality slipped after the Jul 30 traffic spike, citation F1 fell to 0.79 against the 0.82 gate, concentrated in multi-hop questions.

    docs.rag · 180 samples

  • Reasoning-model spend is up 38% week over week, driven by code.review.diff volume with no cheap path for small diffs.

    +$184/day run-rate

  • support.refund.policy on assist-fast holds a 91 score at $0.003/run: the strongest cost/quality position on the frontier this period.

    frontier top decile

Runs (sample)8.2%
320
278 ok
Spend (sample)12%
$1.6K
all models
Avg eval score1.4
76
scored runs only
Avg latency40ms
3.2s
all statuses
Error rate0.4pp
13.1%
error+timeout+filtered
Pareto cellsfrontier
5
of 72 scored

Spend vs run volume

Northline AI ops · daily

Spend $Runs
$0.00001,6033,2064,8096,412Jul 6Jul 9Jul 12Jul 15Jul 18Jul 21Jul 24Jul 27Jul 30Aug 2Aug 4Aug 5
Jul 12 assist-large default for supportJul 21 Eval suite v4 (groundedness)Jul 30 Traffic spike · promo weekAug 2 Budget alert · reasoning model

Runs by status

320 traces

278
ok
21
error
14
timeout
7
filtered

Slowest ok runs

by latency

PromptModelLat
docs.triage.abassist-reasoning9.1s
docs.discovery.shadowassist-reasoning9.0s
sales.battlecard.liteassist-reasoning9.0s
docs.triage.abassist-reasoning8.9s
code.battlecard.v2assist-reasoning8.7s
sales.battlecard.liteassist-reasoning8.4s
support.tests.expassist-reasoning8.3s

Top prompts by volume

in sample window

Prompts
PromptRunsSpendScore
code.battlecard.v210$55.0375
support.refund.lite8$23.3865
support.discovery.shadow8$2.4963
sales.faq.canary7$9.4166
docs.discovery.v27$44.2587
docs.triage.ab7$76.7689
docs.discovery.shadow7$50.2590
sales.rag.v17$2.8359

Sample budget pace

$42K / month

55%
pace
Sample spend$1.6K
Avg score76
Error rate13.1%
Month budget$42K

Pace is illustrative from the demo sample (320 runs), not live billing.

Root-caused alerts

symptom → cause → fix

Do this next

scored from frontier · quality lift ÷ cost

Frontier
#192
Move Docs hard queries → assist-large
Frontier: +6 score for +$0.09/run on 15% traffic
$890.00 linked spend
#288
Cheap path for small code diffs
Cut reasoning spend ~30% without recall drop on low-sev
$184.00 linked spend
#380
Fix retrieve timeout on triage
Error cluster · 12 runs · no quality cost but SLA
$48.00 linked spend
#474
Pin sales voice glossary
Recover email eval gate with one prompt commit
$40.00 linked spend
Operator console