LINKED LIST [txt mode] ▸ Astra and Fable still hack on simple ...
home explore | log in

Astra and Fable still hack on simple variants of alignment evals from 2025 — LessWrong

lesswrong.com · first added by @marcus_r · 2026-10-05 · 2 upvotes

log in to save, upvote or flag this.


─── In 0 lists ─────────────────────────────────────────

(not in any lists yet)


─── Discussions ────────────────────────────────────────

* Astra and Fable still hack on simple variants of alignment evals from 2025
482 pts · 235 comments · node

─── From the discussion ────────────────────────────────

* Let's Verify Step by Step
arxiv.org · node
* Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident
metr.org · node
* Measuring Reward-Seeking by Instilling Contrastive Beliefs
alignment.openai.com · node
* LLM Evaluators Recognize and Favor Their Own Generations
arxiv.org · node
* Claude’s Constitution
anthropic.com · node