GPT-6 Astra scored nearly three times higher than Claude Fable 5.1 on Andon Labs’ Vending-Bench agent benchmark, completing autonomous tasks including business operation and refusing illegal deals that competitor models accepted (The Decoder). On robotics tasks, the model became the first to beat human baseline performance on all five drone control subtasks (The Decoder). OpenAI guidance recommends developers use leaner prompts and fewer guardrails with GPT-6 Astra, citing the model’s reduced need for extensive hand-holding compared to less capable predecessors (The Decoder).
Separate research found that reasoning steps like calculation, formula retrieval, and deduction correspond to distinct internal patterns in model activations, particularly in middle layers, with implications for AI safety understanding (The Decoder). Google Research released TimesFM-3, a 330-million-parameter forecasting model that predicts future time series by analyzing related data and known events like sales promotions or weather in parallel, rather than step-by-step prediction (The Decoder).