Calibrating the proactive-scanning cascade · 350 real Google Chat messages × 3 models · 2026-08-22
The experiment
Every message from three spaces (AI CoE Go-To, AI Watercooler, AI Spend) was classified independently by a
cheap, mid, and strong model with the same rubric: does this author have a live problem an AI CoE advisor could
genuinely help with? The design tension: catch real cries for help — never interrupt chit-chat.
350
messages classified ×3 models
5
unanimous “jump in”
98.3%
cheap↔strong agreement
0
cheap-tier false positives
37
mid-tier-only flags (spam risk)
Tier
Model
Flagged
Rate
Verdict
Cheap
muse-glimmer-30b
5
1.4%
precision 1.00 · recall 0.56 — ideal flagger
Mid
gemini-2.5-flash
46
13.1%
far too liberal — would spam the space
Strong
claude-sonnet-4.5
9
2.6%
consensus anchor — triage tier
✅ Unanimous: good places to jump in
All three models independently said an intake offer would help. Pattern: real pain, in-progress work,
an implicit “who do I ask?”
watercooler · 2023-11-27
Morning! Are we still moving to the OpenAI Enterprise edition now that the dust has settled with Sam Altman? Asking because we are using the API's and see some some charges, wondering if the API calls are more cost effective on Enterprise.
CHEAPmuse-glimmer-30b · Fireworks⚑ FLAG0.68
Author questions cost effectiveness of OpenAI Enterprise versus API usage for current charges.
MIDgemini-2.5-flash⚑ FLAG0.90
The user is asking about the cost-effectiveness of OpenAI Enterprise API calls, which directly relates to inference provider selection and cost optimization.
STRONGclaude-sonnet-4.5⚑ FLAG0.85
Author is asking about cost-effectiveness of different OpenAI pricing tiers (API vs Enterprise) for their current usage, which is a provider selection and cost optimization question.
aicoe-goto · 2025-06-25
We don't know how sub-agents work in the case of Claude code (whether we can override their system prompts).
CHEAPmuse-glimmer-30b · Fireworks⚑ FLAG0.60
Author is seeking guidance on Claude Code sub-agent behavior and system prompt control for an AI-assisted development workflow.
MIDgemini-2.5-flash⚑ FLAG0.80
The user is asking a specific technical question about overriding system prompts for sub-agents in Claude, which suggests a live problem with AI model configuration.
STRONGclaude-sonnet-4.5⚑ FLAG0.75
User has an active question about Claude code sub-agent configuration that could benefit from AI CoE guidance on AI-assisted development tooling.
aicoe-goto · 2025-07-03
I have to jump in an IM meeting, but I would love to hear thoughts about running loops in Agentic AI. I've been struggling to make it work in Cu Chulainn (it always drifts, and the Andon Corder finds a reason to shut it down after a few iterations.)
CHEAPmuse-glimmer-30b · Fireworks⚑ FLAG0.72
Author reports live agentic loop drift and shutdown issue needing troubleshooting.
MIDgemini-2.5-flash⚑ FLAG0.90
The user is struggling with an AI agent's reliability and drift, which is a problem an AI CoE advisor could help diagnose and solve.
STRONGclaude-sonnet-4.5⚑ FLAG0.85
Author is experiencing a live technical problem with agentic AI loop reliability and drift that an advisor could help troubleshoot.
aicoe-goto · 2025-07-11
Hi Team, do we have some sort of documentation / guiding principles how to evaluate RAGs? like which where to get certain LLM keys / where to run the evaluation in infra level and how to minimize cost in the context of XO
CHEAPmuse-glimmer-30b · Fireworks⚑ FLAG0.70
Author is seeking guidance on RAG evaluation, LLM key access, infra placement and cost minimization, indicating a live AI quality/cost problem.
MIDgemini-2.5-flash⚑ FLAG0.90
The user is asking for guidance on evaluating RAGs, including infrastructure, cost minimization, and LLM key management, which are all areas an AI CoE advisor could help with.
STRONGclaude-sonnet-4.5⚑ FLAG0.85
Author is asking for guidance on RAG evaluation methodology, infrastructure setup, and cost optimization, which are actionable AI CoE topics.
aicoe-goto · 2025-07-11
for the LLM Keys, I'll reword that as API keys, for example we have a RAG pipeline and at the moment we're using open AI, ideally we would want to compare it with let's say using claude, grok etc
CHEAPmuse-glimmer-30b · Fireworks⚑ FLAG0.75
Author wants to compare LLM providers for an existing RAG pipeline.
MIDgemini-2.5-flash⚑ FLAG0.90
The user wants to compare different LLMs (OpenAI, Claude, Grok) for their RAG pipeline, which is a direct model selection problem.
STRONGclaude-sonnet-4.5⚑ FLAG0.95
User wants to compare different LLM providers (OpenAI vs Claude vs Grok) for their RAG pipeline, which is a clear model/provider selection problem.
No examples in this date range — try a wider window.
🤫 The cheap model’s misses
Mid + strong flagged these; glimmer-30b stayed quiet. Every miss is short and context-poor — the case
for classifying thread windows instead of isolated messages. Note the asymmetry: a miss costs nothing
(explicit invocation still works); a false positive spends trust.
watercooler · 2023-11-24
Very good point; fact retrieval is a very specific Gen AI task/capability out of many others, like text summarisation, translation, sentiment analysis etc (Chat GPT gives me a list of 18 capability categories).
Do we have insights on what platform is best for each of these? That table will change with every new version, but could be very useful when looking for applications and selecting tools to use?
CHEAPmuse-glimmer-30b · Fireworkspass0.78
General discussion requesting a comparative platform table with no specific live cost, quality or workflow problem stated.
MIDgemini-2.5-flash⚑ FLAG0.90
The user is asking for insights on which platform/model is best for various Gen AI tasks, which directly relates to model selection.
STRONGclaude-sonnet-4.5⚑ FLAG0.85
Author is asking for guidance on selecting the best platform/model for different AI capabilities, which is a model selection problem the AI CoE could help with.
aicoe-goto · 2025-07-10
Cursor be heavy on my M4 Pro
CHEAPmuse-glimmer-30b · Fireworkspass0.75
Vague performance comment with no clear AI cost/quality or workflow help request.
MIDgemini-2.5-flash⚑ FLAG0.80
The message describes a performance issue with an AI-assisted development tool (Cursor) which falls under AI-assisted software workflows.
STRONGclaude-sonnet-4.5⚑ FLAG0.75
User is reporting a performance issue with Cursor (AI-assisted IDE), which could involve configuration or workflow optimization help.
watercooler · 2023-11-29
Which GPT is best for OCR?
CHEAPmuse-glimmer-30b · Fireworkspass0.68
Casual watercooler query with no indication of a live cost or quality problem.
MIDgemini-2.5-flash⚑ FLAG0.90
The user is asking for advice on which GPT model to use for a specific task (OCR), which is a direct model selection problem.
STRONGclaude-sonnet-4.5⚑ FLAG0.85
User is asking for help choosing the right model for an OCR task, which is a clear model selection question.
aicoe-goto · 2025-07-11
Hello Team,
Any recommendation for prompt. tracing tool for agents built using Langgraph ?
My best find is : Self hosted LangFuse.
CHEAPmuse-glimmer-30b · Fireworkspass0.85
General tool recommendation request, not a live cost/quality/reliability problem for AI CoE intake.
MIDgemini-2.5-flash⚑ FLAG0.80
The user is asking for recommendations on a prompt tracing tool for AI agents, which is a specific AI-related problem where an advisor could offer guidance.
STRONGclaude-sonnet-4.5⚑ FLAG0.85
Author is seeking recommendations for tooling to trace/monitor LangGraph agents, which relates to AI workflow quality and observability where an advisor could provide guidance.
No examples in this date range — try a wider window.
🚫 Mid-tier-only flags — the spam the cascade prevents
37 messages flagged by gemini-2.5-flash but rejected by both other models: AI-adjacent enthusiasm,
jokes, and product ideas mistaken for live problems. If this model were the flagger, ~1 in 8 messages would
produce a bot card.
watercooler · 2023-11-10
Product request - More shortcuts like superhuman has
CHEAPmuse-glimmer-30b · Fireworkspass0.95
Message is a generic product feature request with no AI cost, quality, or workflow problem.
MIDgemini-2.5-flash⚑ FLAG0.80
The user is requesting AI-assisted software workflows similar to Superhuman, which aligns with the software_factory category.
STRONGclaude-sonnet-4.5pass0.95
This is a product feature request about keyboard shortcuts, not an AI cost or quality problem requiring CoE advisor help.
watercooler · 2023-11-12
Should be 1m lines of code *deleted* with ai
CHEAPmuse-glimmer-30b · Fireworkspass0.85
Casual watercooler comment with no request for help or described problem.
MIDgemini-2.5-flash⚑ FLAG0.80
The message suggests using AI for code deletion, which relates to AI-assisted development workflows and potentially code quality/cost.
STRONGclaude-sonnet-4.5pass0.85
This is a casual comment or opinion about AI's potential impact on code, not a request for help with a specific problem.
watercooler · 2023-11-12
Hmm, has me thinking, we should do a 5k rewite with AI. I can volunteer Aurea platform.
CHEAPmuse-glimmer-30b · Fireworkspass0.85
Casual suggestion with no specific AI cost, quality or reliability problem described.
MIDgemini-2.5-flash⚑ FLAG0.80
The message proposes an AI-assisted rewrite of a 5k line codebase, which falls under AI-assisted development workflows.
STRONGclaude-sonnet-4.5pass0.85
This is a casual suggestion or idea being floated in conversation, not a live problem requiring AI CoE advisor help.
watercooler · 2023-11-13
When I finished conversational prompting GPT lied to me and said it could not generate images 😱
CHEAPmuse-glimmer-30b · Fireworkspass0.85
Casual watercooler vent about model behavior with no request for CoE assistance.
MIDgemini-2.5-flash⚑ FLAG0.80
The user is experiencing an issue with a conversational AI model's capabilities, which could involve model selection or understanding its limitations.
STRONGclaude-sonnet-4.5pass0.85
This is a casual observation about GPT's behavior shared in a watercooler chat, not a request for help with a cost or quality problem.
watercooler · 2023-11-14
Just off the back of 2 RFP completions with a couple of thousand questions, an Enterprise Custom GPT RFP Responder trained on our product docs and existing responses will be an epic timesaver.
CHEAPmuse-glimmer-30b · Fireworkspass0.85
Statement of an idea with no request for help or live problem described.
MIDgemini-2.5-flash⚑ FLAG0.90
The user is looking for an AI-assisted workflow to improve RFP response efficiency, which falls under AI-assisted software workflows.
STRONGclaude-sonnet-4.5pass0.85
This is a general observation about potential future value of an AI tool, not a request for help with an active problem.
watercooler · 2023-11-14
It should reduce the time to complete a large RFP by at least 50%, maybe nearer to 80% when well trained.
CHEAPmuse-glimmer-30b · Fireworkspass0.85
Message is a general expectation statement with no live problem or request for help.
MIDgemini-2.5-flash⚑ FLAG0.80
The message describes a potential improvement in workflow efficiency (RFP completion time) which could be achieved through AI-assisted software workflows.
STRONGclaude-sonnet-4.5pass0.85
This is a general statement about potential time savings, not a request for help with a specific problem.
watercooler · 2023-11-14
I just shared a new folder with you that contains our 2degrees OCS RFP response and Solution Description. I also added our Information Security Policy and Incident Management Manual. The information in those documents should cover at least 50% of the questions we are typically asked for in a medium size RFP. Also a mix of PDFs, .docx and .xlsx.
Thanks for helping and very interested to see how this turns out. We will need to include the definition of compliance in the training. And nuancing the 'compliance grade' is quite a human skill - until now, maybe !
CHEAPmuse-glimmer-30b · Fireworkspass0.85
Message is an FYI about shared RFP documents and training plans with no specific AI cost/quality problem requested.
MIDgemini-2.5-flash⚑ FLAG0.80
The message discusses using AI to assist with RFP responses, specifically around compliance definitions and grading, which aligns with AI-assisted software workflows.
STRONGclaude-sonnet-4.5pass0.85
This is an FYI message about sharing documents for an RFP response, not a request for help with an AI cost or quality problem.
watercooler · 2023-11-14
Here's some demo videos for responsive (fka rfpio). They are using the OpenAI APIs. While we can create a GPT with all our knowledge and get back answers to questions we type in, this product also provides workflow for identifying questions in docs, prefilling responses and assigning answering responsibility for questions or sections to different people.
Now that we're doing 4 RFPs this month, this type of automation would be valuable.
https://play.goconsensus.com/sac9a97ad
Also the goconsensus tool is interesting, where I'm sure they are seeing metrics about who is watching the content and have probably customized what they sent out in response to the qualification call I did.
CHEAPmuse-glimmer-30b · Fireworkspass0.85
Shares demo videos and notes potential value but does not request help or describe a live problem.
MIDgemini-2.5-flash⚑ FLAG0.90
The message describes a need for AI-assisted workflow automation for RFP responses, which aligns with AI-assisted software workflows.
STRONGclaude-sonnet-4.5pass0.85
This is a product demo share and general discussion about RFP automation tools, not a request for help with an active AI cost or quality problem.
No examples in this date range — try a wider window.
Recommendations (now in PRD v0.6 §8.9)
Cheap flagger (muse-glimmer-30b) → strong triage; never a liberal mid-tier as flagger.
Classify thread windows (parent + recent replies), not isolated messages.
Flag only at cheap-confidence ≥ 0.7, then require strong-triage agreement before posting.
Precision >> recall for proactive suggestions; misses are free, false positives spend trust.
Volume is low (~1.4–2.6% of messages) — per-person-per-day cap stays as cheap insurance.
Demand skews to model selection / RAG evaluation / agentic workflows — prioritize those playbooks.