We can't find the internet
Attempting to reconnect
Something went wrong!
Hang in there while we get back on track
Top Stories
Why it matters: Model competition is becoming a contest over reliable, affordable execution—not just headline scores.
Gemini 3.7 Flash makes rapid, cheap iteration the headline. Google introduced its “most intelligent workhorse” three weeks after 3.6 and reports gains from 34.4% to 43.6% on FrontierCode, 49.0% to 65.3% on DeepSWE, 1538 to 1588 WebDev Arena Elo, and 17.0% to 30.4% on AutomationBench. The introductory API price is $0.75/$3.75 per million input/output tokens through 2026, rising to $1.50/$7.50 in 2027; access spans developers, enterprises, and individuals through Spark for Google AI Pro and Ultra subscribers.
DeepSeek is shipping a model and a programmable harness layer. V4 Pro adds low/high/max reasoning effort, native OpenAI Responses API support optimized for Codex, and app/API access. An accompanying release thread identifies V4 Pro 0813 as an MIT-licensed open-weight checkpoint on Hugging Face; Harness v0.1 is also MIT-licensed and makes models, tools, sessions, sandboxes, loops, orchestration, and UI plugins. DeepSeek says new off-peak API rates will be 50% below peak, effective August 16, adding scheduling as another lever for agent economics.
Research & Innovation
Why it matters: The hard production problems are state retention, instruction overhead, and whether evaluations generalize.
Context compaction can erase operating constraints. A COMPINT evaluation summary says current compactors retain only 17% of standing rules, silently dropping session instructions such as “do not delete any emails until I confirm”; compacted runs can be worse than running without compaction. An SC-aware extractor recovered more than 90% retention without changing the model or compactor.
Skill libraries are not free guidance. A Microsoft-and-colleagues paper summary attributes 307 agent failures to loaded skills—125 functional failures and 182 efficiency regressions. Seemingly relevant skills sometimes caused agents to omit or misimplement requirements; excessive verification accounted for 67 cost regressions and heavy implementation pipelines for 30.
Agent leaderboards may rank specialization. A four-facet Generalizability Theory analysis across TheAgentCompany, tau-squared-bench, and AppWorld finds the agent effect explains under 3% of variance while agent-by-task interaction explains 7–23%; on the hardest quartile, reliability falls from 0.752 to 0, and per-family rankings invert.
Products & Launches
Why it matters: The execution layer is becoming a product surface, from inference speed to prebuilt environments and hands-off orchestration.
OpenAI’s Ultrafast mode, powered by Cerebras, promises up to 750 tokens per second—14× faster than standard GPT-5.6 Sol. It starts with a select API customer group and targets real-time voice, support, commerce, coding, financial research, and security response.
Cursor says prebuilt “builds” cut cloud-agent startup time threefold at no additional cost; failed builds never go live, and customers report start times falling from minutes to seconds.
NAC brings long-running delegation into an open harness. Launched with a beta expanded Open Models API, it was used daily by its research team since April for asynchronous, hands-off work and powered a significant portion of recent pre-training, post-training, and data-pipeline code before opening to everyone.
Industry Moves
Why it matters: Capital and infrastructure are following agents into governed data systems, observability, and national-scale compute.
Databricks says it crossed a $7B revenue run-rate, up more than 80% year over year in Q2, and raised $5B to invest in Lakebase, its serverless Postgres for AI agents; Genie, its business-data AI coworkers; and Unity AI Gateway for multi-AI governance and cost control.
Together AI and Larsen & Toubro are building a 10,000-Nvidia-B300 “AI Factory” in India, aimed at open-source inference, fine-tuning, and training at scale.
Arize entered a definitive agreement to be acquired by Dynatrace. Arize’s founder frames the deal around the convergence of software and agents: tools and prompts mix code, while software logs and traces help debug AI systems.
Quick Takes
Why it matters: Open and specialized releases keep widening the set of deployable alternatives.
- GLM-5.3: Z.ai positions the model for coding and cyber defense after post-training on a 743B base; it is available through GLM Coding Plan and ZCode, with API access and open weights staged after safety evaluations.
- dots3-note: Dots Studio’s preview is a 280B MoE with 16B active parameters, 512K context, multimodal input, and TEMPO for long-horizon agent training; vLLM says it is Apache 2.0 with day-one vLLM support.
- LlamaExtract Agentic Plus: LlamaIndex describes a document-extraction model-plus-harness engine; its release claims 95.6% value accuracy at less than a third of the closest peer’s cost.
- MiniMax-H3: Arena places it first overall in Video Edit Arena at 1,390 points, 32 points ahead of the next two models.


