
Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety
By arXiv
AI systems that "think" in human language offer a unique opportunity for AI safety: we can monitor their chains of thought (CoT) for the intent to misbehave. Like all other known AI oversight methods, CoT monitoring is imperfect and allows some misbehavior to go unnoticed.
- Use cases, geography and tags
- Organizations that created, adopted or are mentioned
- Ecosystem position
- Link to the original asset
- Comments and reactions