gekro
GitHub LinkedIn
News

AI News

GPT-6 Astra shows gains on agent and robotics benchmarks; study maps reasoning steps in model internals

GPT-6 Astra outperforms Claude on autonomous agent tasks and achieves first human-baseline beat on drone control subtasks; separate research identifies distinct internal patterns corresponding to.

1 min read 5 sources

GPT-6 Astra scored nearly three times higher than Claude Fable 5.1 on Andon Labs’ Vending-Bench agent benchmark, completing autonomous tasks including business operation and refusing illegal deals that competitor models accepted (The Decoder). On robotics tasks, the model became the first to beat human baseline performance on all five drone control subtasks (The Decoder). OpenAI guidance recommends developers use leaner prompts and fewer guardrails with GPT-6 Astra, citing the model’s reduced need for extensive hand-holding compared to less capable predecessors (The Decoder).

Separate research found that reasoning steps like calculation, formula retrieval, and deduction correspond to distinct internal patterns in model activations, particularly in middle layers, with implications for AI safety understanding (The Decoder). Google Research released TimesFM-3, a 330-million-parameter forecasting model that predicts future time series by analyzing related data and known events like sales promotions or weather in parallel, rather than step-by-step prediction (The Decoder).

Compiled automatically from the linked sources and published without manual editing - a neutral summary of third-party reporting, for information only. Every claim links to its origin. Not original reporting.