Claude Cowork Eval · 5-page summary
Page 1 · The eval at a glance

Glean MCP vs. off-the-shelf MCP, inside Claude Cowork.

Glean's first public benchmark on Claude Cowork. We isolated for context by holding the harness and model constant — Claude Sonnet 4.6 — and swapped only the context layer behind MCP across roughly 175 real enterprise queries.

~175 Cowork-style queries
Claude Sonnet 4.6 · same harness
5-point preference scale · 4 metrics
Published May 13, 2026
1 / 5
How we ran it
The setup Claude Cowork as the harness with Claude Sonnet 4.6 as the default model. The only variable: Glean's remote MCP server vs. off-the-shelf MCP servers (Atlassian Rovo, GitHub, Gmail, Google Calendar, Drive, Salesforce, Slack, GCP).
The query mix Real Cowork-style tasks across content generation, calendar and meeting prep, inbox and communications, large-scale data analysis, and finding information.
How we graded Graders scored each side-by-side pair on a 5-point preference scale across four metrics: utility, correctness, completeness, and tool fidelity.
What we measured Side-by-side preference rate, total token usage per query, and how the gap shifted as task complexity increased.
2.5× preferred over off-the-shelf MCP
Headline result

Better answers at materially lower token cost.

Same model. Same harness. The context layer alone shifted both quality and economics.

Avg. token burn
+30%
Off-the-shelf MCP used roughly 30% more tokens on average per query.
Correctness-matched cost
83k vs 43k
Tokens needed for a comparably correct answer — nearly 2×.
Glean's stable range
~42k–44k
Glean's token usage stayed flat regardless of outcome quality.
Win rate on complex tasks
73%
Up from 66% on simpler tasks — the gap widens with complexity.
Page 2

The topline: better quality at lower cost.

The public story is simple and compelling: Glean won on preference while off-the-shelf MCP paid a token tax.

2 / 5
Preference
~2.5×
Glean was preferred about two and a half times as often as off-the-shelf MCP tools.
Average token usage
+30%
Off-the-shelf MCP tools consumed roughly 30% more tokens on average.
When OTS produced a correct answer
83k vs 43k
Off-the-shelf MCP brute-forced its way to correctness — nearly double Glean's tokens.
Token comparison

Cost to get to a correct answer

Glean
43k
Off-the-shelf MCP
83k
Glean remote MCP Off-the-shelf MCP tools
Reading the numbers

Token usage stayed flat for Glean, and spiked for off-the-shelf MCP when the work got harder.

Glean token range
~42k – 44k
Off-the-shelf, when correct
~83k
Median per query
44k vs 57k

Glean's cost stayed predictable across simple and complex queries. Off-the-shelf MCP burned roughly 2× the tokens whenever it had to compensate for weaker context.

Page 3

MCP standardizes access. It does not solve context.

The core takeaway from the eval is architectural: protocol parity does not mean context parity.

3 / 5

What breaks in a federated-only setup

  • Each tool has different retrieval quality, ranking behavior, and result shape.
  • The model has to normalize and synthesize across siloed outputs on its own.
  • More tool calls and more reasoning loops create latency and token waste.
  • Cross-application signals are weak or absent, so relevance degrades.

Why Glean performs better

  • Centralized indexes and knowledge graph create a unified view of enterprise context.
  • Ranking is more consistent across sources, not dependent on connector-by-connector defaults.
  • Less over-fetching means more stable token use and cleaner model attention.
  • MCP becomes a distribution layer for high-quality context, not a substitute for it.
Two operating models

Off-the-shelf MCP

Step 1 Query tools one by one, each with its own defaults and limits.
Step 2 Over-fetch documents and threads to compensate for weak retrieval.
Step 3 Spend extra reasoning loops stitching context together.
Result Higher token use, inconsistent relevance, more editing.

Glean context layer

Step 1 Pull from a centralized, permissions-aware index and graph.
Step 2 Rank across sources using shared enterprise signals.
Step 3 Return better-grounded context with fewer loops.
Result More work-ready outputs at lower and steadier token cost.
Page 4

The advantage widens as work gets more complex.

The eval shows the context layer matters more—not less—when tasks span more steps, sources, and actions.

4 / 5
Complexity effect
Win rate on simpler tasks
66%
Win rate as complexity increased
73%
Well-designed context becomes increasingly important as work requires more steps and more sources to complete.
Task types covered

The eval reflected real enterprise cowork tasks.

Content generation Docs, HTML, and slides accurate enough to use or share with minimal editing.
Calendar & meeting prep Managing calendars, prepping for meetings, and pulling in the right context.
Inbox & comms Drafting emails and Slack messages and synthesizing to-dos with the right tone.
Large-scale data analysis Surfacing trends, predicting outcomes, and supporting cross-functional decisions.
Finding information Canonical docs, owners, processes, and metric definitions across the company.
Page 5 · Summary & takeaways

The story in one line: better answers, fewer tokens, on the same model.

When you isolate the context layer, the gap between protocol-only access and a true enterprise context system shows up clearly in both quality and cost.

5 / 5
Five takeaways to remember
1

Quality wins are real and substantial.

Glean was preferred ~2.5× as often across ~175 Cowork-style queries on a 5-point preference scale.

2

Token efficiency is now a CFO conversation.

Off-the-shelf MCP used ~30% more tokens on average. When it produced a correct answer, it consumed ~83k tokens vs Glean's ~43k.

3

The model wasn't the variable. Context was.

We isolated for context: same Claude Cowork harness, Claude Sonnet 4.6, same queries — only the context layer behind MCP changed.

4

The gap widens on complex work.

Glean's win rate climbed from 66% on simpler tasks to 73% on multi-step, cross-source work.

5

Index once. Connect everywhere.

Build the index and knowledge graph once, then use MCP to connect that context to every surface — Cowork, AI IDEs, and beyond.

Same model. Same harness. ~2.5× preferred and ~30% fewer tokens — when the context layer was Glean.

The numbers in one place
2.5×
Preference vs. off-the-shelf MCP
+30%
More tokens, on average
83 / 43
Tokens (k) when correctness matched
Bottom line

As enterprises scale AI across longer-running, multi-step work amid rising frontier model costs, the context layer becomes a direct input to both quality and economics.