0:00 / 3:59
Chapters
Sources
DAILY ROUNDUP
OpenAI And Cerebras Launch Ultrafast Mode, GPT-5.6 Sol At 750 Tokens/Sec
calendar_today Date:
schedule Duration: 3:59
visibility 2 Views
OpenAI and Cerebras launch Ultrafast mode for GPT-5.6 Sol at 750 tokens/sec, DeepSeek ships V4-Pro with flexible reasoning effort, Google releases Gemini 3.7 Flash, and MiniMax open-sources a production-ready music model.
- 01. GPT-5.6 Sol on Ultrafast runs 14x faster than Standard, completing Humanity's Last Exam in 11 hours versus 78 hours for Claude Fable 5
- 02. DeepSeek-V4-Pro adds flexible reasoning effort (low/high/max) and native OpenAI Responses API support with one-click Codex setup
- 03. Gemini 3.7 Flash ships just three weeks after 3.6 Flash with real coding benchmark gains and a 50% introductory price cut
- 04. MiniMax Music 3 generates complete songs up to five minutes long, open-weighted and free for commercial use under Apache 2.0
Today's AI news: OpenAI and Cerebras launched Ultrafast mode, running the full GPT-5.6 Sol model at up to 750 tokens per second, roughly 14x faster than standard processing. DeepSeek launched V4-Pro with major agent upgrades and flexible reasoning effort. Google released Gemini 3.7 Flash, its most intelligent workhorse model yet for coding and agents. And MiniMax open-sourced Music 3, a production-ready music model.
Chapters:
0:00 Today's AI News
0:30 GPT-5.6 Sol Ultrafast
1:37 DeepSeek V4-Pro
2:24 Gemini 3.7 Flash
3:18 MiniMax Music3
Blend Roundup 2026-08-14
https://x.com/OpenAI/status/2087947721936359705
https://x.com/cerebras/status/2087948820906950719
https://x.com/cerebras/status/2087961128869748856
https://x.com/deepseek_ai/status/2087864585504305397
https://x.com/Google/status/2087948901265354817
https://x.com/antigravity/status/2087953243552743525
https://x.com/MiniMax_AI/status/2087934657354678421
[quick] OpenAI and Cerebras have launched Ultrafast mode, running GPT-5.6 Sol at 750 tokens per second, DeepSeek launched V4-Pro with major agent upgrades, Google released Gemini 3.7 Flash for coding and agents, and MiniMax open-sourced a new production-ready music model. [excited] Here's today's AI news.
[quick] OpenAI and Cerebras have launched Ultrafast mode for GPT-5.6 Sol, a new tier that runs the full model at up to 750 tokens per second, roughly 14 times faster than standard processing, without any drop in the model's underlying intelligence.
On Humanity's Last Exam, a 2,500-question benchmark pitched at PhD-level difficulty, GPT-5.6 Sol on Ultrafast worked through the whole set in just over 11 hours, versus roughly 78 hours for Claude Fable 5 on the same test.
In a head-to-head build test, generating a financial dashboard, Ultrafast finished in under two minutes against twelve minutes on Standard, with the same result both times.
It's launching first to a select group of OpenAI API customers, with wider access coming as capacity grows, and it makes raw response speed a genuinely new lever labs are competing on, not just intelligence.
[quick] DeepSeek has launched DeepSeek-V4-Pro, its new flagship model, bringing major upgrades to how it handles agent workflows in production.
Both V4-Pro and its smaller sibling V4-Flash now support flexible reasoning effort, so you can dial it low for simple tasks, high for everyday agent workflows, or max for genuinely hard problems, rather than getting one fixed setting.
V4-Pro runs at 1.6 trillion parameters with 49 billion active, supports a million-token context window, and adds native support for OpenAI's Responses API, including one-click setup for Codex.
It's a serious production-focused update from one of the few labs still shipping frontier open models at this scale.
[quick] Google has released Gemini 3.7 Flash, which it's calling its most intelligent workhorse model yet for coding and agents, arriving just three weeks after Gemini 3.6 Flash.
It shows real gains on coding benchmarks, producing more production-ready code and handling debugging and multi-step planning more reliably, and it's already live in Google Antigravity for anyone wanting to try it in an actual coding workflow.
Google is pricing it at 75 cents per million input tokens and three seventy-five per million output tokens through the end of the year, roughly half what 3.6 Flash launched at.
Shipping a meaningfully better model just three weeks after the last one is an unusually fast pace even by this year's standards.
[quick] MiniMax has open-sourced Music 3, a new production-ready music model that generates complete songs up to five minutes long from a description and optional lyrics.
It combines an eight-billion-parameter model for overall musical structure with a smaller local model handling frame-level detail, producing full 32-kilohertz stereo audio with expressive vocals and arrangements that actually evolve over the length of the track.
The weights are open and free to use commercially under the Apache license, so anyone can run it locally or plug it into tools like ComfyUI rather than relying on a hosted API.
For a field where most of the serious music generators have stayed closed, that's a genuinely different approach.