事实:事件:Last week, OpenAI’s new GPT-6 Astra was released with a big fanfare. I used it over the past couple of days, and it’s an exceptionally good model, likely the best I’ve used as of this writing. But what, exactly, has it improved, and how? 1.1 Astra benchmarks Astra is the best model I’ve used so far, and it’s disproportionately good at 3D rendering and animation tasks (relative to other models). With that, I mean that while it leapfrogs its GPT-5.6 predecessor in practically all categories (writing, math, coding, and more), it especially does so when it comes to graphical demos. We can see this also reflected in the benchmarks. For instance, GPT-6 Astra is really good at math and coding, as shown below. Figure 1: Selection of three popular coding benchmarks and one challenging math benchmark. More benchmarks are shared on the Astra release blog: https://openai.com/index/gpt-6-astra/ One of the highlights (not shown in the figure) is that Astra also achieves 99.9% on the ARC-AGI-3 benchmark (GPT-5.6 Sol only 7.8%), which measures a mix of solving logic puzzles and generalization.|影响:适合需要评估模型能力的团队,用于决定是否启动内部测试。|行动:建议用本团队的图形、写作、数学和编码任务进行对比测试。|风险:建议以能否复现实测表现为边界,不能复核就不用该结果。
正在加载入行工作台