LINKED LIST [txt mode] ▸ Beyond Benchmarks: Disagreement Among...
home explore | log in

Beyond Benchmarks: Disagreement Among Frontier LLMs on Real-World Fact-Checks (v1.1)

lenz.io · first added by @hnl · 2026-09-25 · 2 upvotes

log in to save, upvote or flag this.


─── In 1 list ──────────────────────────────────────────

* Issue #795 by @hnl [data]

─── Discussions ────────────────────────────────────────

* Disagreement among frontier LLMs on real-world fact-checks
505 pts · 347 comments · node

─── From the discussion ────────────────────────────────

* Beyond Benchmarks: Disagreement Among Frontier LLMs on Real-World Fact-Checks
doi.org · node
* Datasette
lite.datasette.io · node
* lenz_llm_disagreement.csv
docs.google.com · node

─── Also saved alongside this ──────────────────────────

* You're Not Burnt Out. You're Existentially Starving. - Neil Thanedar
3 upvotes · 2026-02-27
* I think Anthropic and OpenAI have found product-market fit
3 upvotes · 2026-07-18
* Using AI to write better code more slowly
3 upvotes · 2026-07-18
* I’m tired of talking to AI
2 upvotes · 2026-09-25