ALPHAZETARESEARCH & EVIDENCE

Tooling case study

How we used agents to evaluate Jev—and changed our minds about its best use

Published · Research through

Executive summary

We used agents to test Jev on research tasks. The strongest finding was a narrow, practical one: it could help prioritise claims for closer inspection. Its scores could not establish that a claim was correct.

Agents compared approaches and checked results against sources, narrowing the recommendation as evidence accumulated. These remain internal, model-reviewed findings—not an independent human benchmark. Research through 24 September 2026; Jev 1.13.0.

Open a finding to explore it.