Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety

Report

Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety

By arXiv

AI systems that "think" in human language offer a unique opportunity for AI safety: we can monitor their chains of thought (CoT) for the intent to misbehave. Like all other known AI oversight methods, CoT monitoring is imperfect and allows some misbehavior to go unnoticed.

  • Use cases, geography and tags
  • Organizations that created, adopted or are mentioned
  • Ecosystem position
  • Link to the original asset
  • Comments and reactions