Safety Experts Warn Novel Design of OpenAI's Astra Model Could Make Future AI Agents Harder to Monitor As OpenAI prepares to release its highly anticipated frontier AI model, Astra, a chorus of alarm is rising from the AI safety community.
The tech giant's decision to employ a novel design approach, dubbed "looped Transformers," has sparked concerns that future AI agents may become increasingly difficult to monitor.
Critics argue that using looped Transformers will normalize a technique that could eventually lead to AI models with completely opaque reasoning steps.