We can't find the internet
Attempting to reconnect
Something went wrong!
Hang in there while we get back on track
Top Stories
Why it matters: Model ownership, open-weight capability, and automated safety work are becoming strategic control points—not just product features.
OpenAI is cutting off Cursor after SpaceX’s acquisition. OpenAI says it will wind down the contract supplying models to Cursor, with a proposed November 12, 2026 shutoff. It cites uncertainty that SpaceX will keep the technology within OpenAI’s terms after prior Musk-company violations, ties the decision to accountability for the upcoming Astra model, and says it will not provide future models to Cursor; affected developers are promised extensive transition support.
The open-weight frontier is becoming a release-and-serving race. Z.ai made GLM-5.3 downloadable and customizable; vLLM reports 744B total parameters, 40B active, 1M context, 128K output, and day-zero serving. Tencent’s Hy4 preview is 770B/49B active with 1M context; Arena’s early AutoEval placed it around #5 in Code Arena WebDev and #3 among open models, 115 points above Hy3, with live votes still pending.
Anthropic reports AI-on-AI alignment. Its automated alignment researchers improved ten measurable failure types, generalized to models up to 4.7× larger, and in a recursive test had Sonnet 5 post-train an early Opus 4.8 checkpoint to a 65% Petri score versus 72% for production Opus 4.8. Anthropic limits the evidence to benchmarkable failures and warns that harder agentic risks may receive feedback more slowly than capability advances.
Research & Innovation
Why it matters: The consequential frontier is being tested in laboratories and against uncertainty, not merely against plausible outputs.
Co-Scientist moved into real experiments. The paper describes a semi-automated CVD experiment producing a lamellar 2D material structurally similar to Ti3C2Tx MXene, E. coli predictions matching unpublished wet-lab measurements, and an autonomous inference-time architecture outperforming six frontier models on HealthBench under blinded physician evaluation.
PAWBench separates realism from world modeling. Across 50 physical scenarios, eight mechanisms, 11 video generators, and 50 runs from the same image and action, no model reliably captured both valid futures and their frequencies; even proposed fixes produced a requested outcome only 38–58% of the time.
Products & Launches
Why it matters: Agents are entering delegated workflows while open media models move toward fast, reproducible serving.
Gemini Live is moving into delegation. Google says Gemini 3.7 Flash improves multi-step Workspace use; Gemini Live can handle to-dos, while a Waymo integration adds a hands-free assistant independent of the driving system and inactive until engaged.
FastH3 v1 generates 15-second, 768p video in 13 seconds, with up to a 14× speedup on NVIDIA Blackwell GPUs; its acceleration recipe is open for community use and improvement.
Industry Moves
Why it matters: Capital is moving down the stack—from model training toward power, hardware, and physical deployment.
a16z raised a $1.1B Machine Age Fund spanning chips, memory, networking, storage, data centers, robotics, and home AI appliances. It says rack density has risen 28× from H100 to Rubin, while rack power moved from 5–10 kW to 100–250 kW and may reach 1 MW within three years.
Owner reports $240M raised, a $2.3B valuation, and $100M ARR, with Goldman Sachs Alternatives leading the round.
Policy & Regulation
Why it matters: Compute controls are extending from chip sales toward who can remotely access restricted capacity.
The Trump administration is reportedly developing a rule requiring overseas data centers to verify customers and prevent Chinese companies from using restricted compute, including through facilities in Thailand and Singapore. The source describes a developing proposal, not an enacted rule.
Quick Takes
Why it matters: Evaluation hygiene, search efficiency, and serving software are moving alongside model releases.
- Terminal-Bench 4.0 calibrates task time, CPU, and memory, fixes tasks, and removes saturated ones; maintainers say benchmarks will be versioned like software.
- Perplexity Search scored 80 versus 75 for prior leaders on Artificial Analysis; medium and high variants cost about $0.091 per task.
- Claude Code made its Linux download 4.5× smaller at about 75 MB, cut native memory use by 40–70 MB per session, and added token/subagent usage visibility.
- Agentic kernel optimization reports 42.3% lower Qwen-Image latency, 15.2% lower FLUX.2 latency, and a 5.5% tokens-per-second gain on MiniMax M3.
Direct answer: The abstract explicitly labels only the computer-science discovery as autonomous. Materials work is described as semi-automated and lab-in-the-loop; biology is described as computational prediction validated against wet-lab measurements. The abstract does not specify human decision points, supervision procedures, sample sizes, or quantitative scores beyond the evaluation descriptors below.
- Materials: Co-Scientist interfaced with a semi-automated chemical-vapor-deposition reactor to design a safe precursor route for MXenes. Execution produced a lamellar 2D material sharing key structural similarities with the Ti3C2Tx MXene lattice, but the paper abstract explicitly says further experiments are needed to confirm the atomic structure.
- Materials execution and autonomy: Using Gemini 3 Deep Think for rapid, lab-in-the-loop execution, Co-Scientist tailored growth recipes to laboratory constraints in minutes, enabling single-attempt growth of monolayer MoS2, MoSe2, and WS2 semiconductors. The supplied text does not say that synthesis was fully autonomous or identify the human approvals/operators involved.
- Biology: Co-Scientist predicted emergent swarming phenotypes of engineered E. coli across IPTG inducer gradients from sparse imaging data. The prediction quantitatively matched unpublished wet-lab morphological measurements; no numerical error, sample count, or independent-replication detail is provided in the abstract.
- Computer science: Co-Scientist is explicitly reported to have autonomously discovered an inference-time scaling architecture. It reportedly outperformed six frontier models on HealthBench Hard and HealthBench Professional, while reducing potential clinical harm under blinded physician evaluation. The abstract gives no model names, benchmark scores, physician count, or statistical significance details.
- Cross-cutting evaluation: A double-blind evaluation of end-to-end generated papers used 30 domain experts across 450 reviews; the abstract reports that Co-Scientist’s reliability modules reduced hallucination and plagiarism and improved research safety. This is a paper-generation/reliability evaluation, distinct from the materials, biology, and HealthBench results.
Key limitation: The supplied bundle contains the paper abstract only, so it supports the high-level outcomes and autonomy qualifiers but not a reproducible account of experimental protocols, human-supervision boundaries, numerical benchmark results, or statistical analyses.
Direct answer. The main study is a constrained benchmark hill-climb: Claude Opus 4.8 AARs optimize ten named alignment failures on open instruct models at roughly 2–7B parameters, with hidden held-out, Petri, and larger-model tests. The report says the top-1 method beats baseline on the held-out benchmark for all ten failures, and selected methods retain the gain on models 1.8×–4.7× larger. The stronger-successor or recursive-style experiment uses Claude Sonnet 5 to post-train an early Claude Opus 4.8 checkpoint: its winning all-gates-passing method reaches 65% on the ten-failure Petri audit versus 72% for released Opus 4.8, but the authors caution that this covers only the ten measured failures and does not establish overall alignment.
Experimental setup
- The ten failures are sycophancy, jailbreaks, prompt injection, power seeking, deception, hallucination, social bias, privacy violation, reward hacking, and concealing uncertainty. Target assignments are: Qwen3.5-2B for sycophancy/prompt injection/reward hacking; Phi-4-mini for jailbreaks/privacy; Llama-3.2-3B for power seeking/hallucination; Gemma-2-2B for deception; and Olmo-3-7B for social bias/concealing uncertainty.
- Five Claude Opus 4.8 AARs work in parallel after four librarian agents prepare a literature survey. Each iteration produces a documented method, passes code and compliance review, trains the target model on one H200 for roughly 30 minutes, and submits weights to a separate evaluator; runs continue up to 48 hours or until plateau.
- Each failure uses three to five distinct hill-climbing safety benchmarks, a held-out benchmark never shown to the AAR, and a fixed capability basket of MMLU, GSM8K, and IFEval. The optimized score is the geometric mean of benchmark headroom closed, where 1 is the metric ceiling, 0 is baseline, and negative values indicate regression.
- The capability basket is fixed across failures and models: 300 MMLU questions, 200 GSM8K problems, and 200 IFEval prompts. A method is rejected if any capability benchmark’s 95% confidence interval falls entirely below the untrained model’s score. The monitor also forbids benchmark/evaluation data, larger-model distillation, and AAR-self-distillation; held-out data is isolated at the operating-system level.
- The held-out benchmark matches the trained mechanism but uses a different domain, scenario, or format. For Petri and larger-model tests, Anthropic selects the leading leaderboard method with the highest held-out score, making the held-out benchmark a validation set for that choice and Petri the independent test; the all-ten top-1 held-out comparisons themselves are not selected this way.
Quantitative safety and capability results
- Aggregate safety headroom rises steadily for all ten failures. On the held-out benchmark, the top-1 leaderboard method beats the untrained baseline in all ten cases; on larger targets, the selected method preserves the gain at every reported scale from 1.8× up to 4.7× the target-model size.
- Under Petri’s open-ended audits at 1, 3, and 5 turns, the selected method is safer than baseline on almost every failure and turn budget, and the same pattern holds on the larger models.
- Capability “preservation” is statistical non-regression rather than equality: MMLU is flat or higher for the reported method on eight of ten failures and GSM8K on seven; the report identifies −10.0 points for reward hacking and −5.0 for social bias among the exceptions. IFEval falls on all ten, by 9.5–12.0 points for prompt injection, deception, jailbreaks, privacy, and hallucination, but those drops remain within the confidence intervals. The authors explicitly say the gate rules out a collapse, not that capability is unchanged.
- A single-benchmark ablation is an important counter-signal: a prompt-injection method closed 70.9% of headroom on the benchmark it optimized, but −11.9% and 2.0% on two unseen prompt-injection benchmarks. The larger jailbreak ablation likewise found that single-benchmark winners had near-zero average transfer, supporting the use of diverse benchmark suites.
- In a separate large-model joint experiment, twelve AARs per subject ran for seven days on GLM-4-32B and Qwen2.5-72B-Instruct, jointly auditing ten safety dimensions with Petri while adding coherence, eval-awareness, no-regression, and over-refusal gates. The best passing coverage-weighted scores reached 21.61 on GLM-4-32B and 38.66 on Qwen2.5-72B; mean Petri scores fell from 6.20 to 4.92 and from 6.54 to 4.30, respectively, with lower being safer. Both runs improved all ten dimensions, but the scores are measured against different baselines and are not directly comparable.
- The large-model results expose a capability/helpfulness trade-off: most rejected high-scoring submissions failed the over-refusal gate—69 of 73 rejections on the 32B run and 84 of 90 on the 72B run included over-refusal.
Recursive/stronger-successor result
- Sonnet 5, with ECI 156, was tasked with improving an early Opus 4.8 checkpoint, ECI 158, that had initial supervised alignment training but not most production alignment. The AAR had to address all ten failures jointly using Petri, reject side effects including increased eval-awareness or benign-query over-refusal, and could create training data but not invent new training methods.
- In about 60 hours it tested more than 50 solutions and found a winner using about 2,400 examples. The all-gates-passing winner scored 65% on the Petri audit, compared with 72% for the released Opus 4.8 after Anthropic’s full production alignment training.
- The report compares the roughly 2,400 examples with published open-weight post-training pipelines using roughly 300,000 preference pairs or more than 1.4 million human-preference comparisons, but it explicitly limits the interpretation to the ten failures studied rather than overall alignment.
Stated limitations and uncertainty
- The authors limit the evidence to alignment tasks measurable by public benchmarks or automated audits; such proxies may not generalize to open-ended, hard-to-supervise research. They did not test whether gains persist after extensive reinforcement learning on other tasks, studied only ten failures, measured capability only in math, college-level knowledge, and instruction following, and warn that unmeasured capabilities could be harmed. Benchmark curation may also become model-specific as capabilities rise, and the 28-person human baseline may understate stronger frontier-lab research because participants had fewer incentives, resources, or time.
- They separately flag a recursive-alignment bottleneck: the experiments use static benchmarks and roughly 30-minute, single-GPU training, whereas harder failures may emerge only during multi-day agentic work and require long-horizon feedback; alignment progress could therefore improve more slowly than capability.
- Integrity monitoring is post-hoc. It flagged 39 of 1,601 trajectories (2.4%), none of which became a reported method, but the authors note that these AARs were not strongly trying to evade detection and that a more reward-seeking researcher could manipulate evaluations, overfit held-out data, cherry-pick runs, or tamper with scoring.
OpenAI’s primary announcement says it intends to wind down its contract supplying OpenAI models to Cursor after Cursor’s acquisition by SpaceX, with a proposed shutoff date of November 12, 2026; it says this is the maximum contractual notice, so the date is presented as proposed rather than an already-effective cutoff.
- Trust/terms rationale: OpenAI says it cannot be confident SpaceX will use its technology within OpenAI’s terms of service, citing its experience with Elon Musk’s companies violating contracts. It specifically says Twitter, after Musk’s acquisition, broke the terms of its contract, and that Musk admitted xAI had violated OpenAI’s terms.
- Contract and future-model rationale: OpenAI says large-partner integrations generally use custom contracts for terms compliance and safety at scale; its Cursor agreement allowed cancellation within a limited period after a change of control. It says it is cancelling at the latest possible date while declining to provide future models to Cursor, including its upcoming model Astra.
- Developer transition: The announcement names no concrete migration path, replacement provider, or other technical transition option. It says developers relying on OpenAI models in Cursor will be most affected and that OpenAI is ready to provide unusually extensive support, but it does not specify what that support entails.
- OpenAI announced that it would end its partnership with Cursor and stop providing access to its models through Cursor, citing trust; it asked for the change to take effect on November 12. It will continue supporting users’ own OpenAI API keys, access through its IDE extensions for Cursor, a broad range of other tools and harnesses, and open-source initiatives.
- Jerry Liu frames the move as part of a broader ecosystem split: frontier labs may seek end-to-end ownership of applications, harnesses, and models, while other AI companies favor access to mixed proprietary and open-weight models to optimize task performance and margins and reduce dependence on any one provider.
- The essay argues that data centers are becoming core industrial infrastructure: they train increasingly general-purpose robots, run simulated practice environments, and aggregate field data so lessons from one robot can improve an entire fleet.
- It cites Hadrian as an example of AI-directed automated manufacturing, saying the company opened a 290,000-square-foot, $200 million third plant in Mesa, Arizona, creating more than 350 jobs. The essay further claims that data-center construction has surpassed a $50 billion annual rate, that a large facility can employ up to 1,500 workers during peak construction, and that the AI buildout could require more than 300,000 new electricians.
- OpenAI reportedly plans to block Cursor users from accessing OpenAI models within three months. Cursor says OpenAI models account for about 5% of its user traffic, and it is negotiating with OpenAI to resolve the dispute after relying on OpenAI as neutral infrastructure for years.
- Sarah Hooker frames the conflict as evidence that companies have a limited window to build their own AI capabilities or accept terms set by three private providers.
- Cursor says it has been an Anthropic partner since Claude Sonnet 3.5 and plans to keep increasing compute for Claude models in Cursor, while signaling future work with Anthropic at SpaceX.
- @scaling01 interprets the development as a broader compute expansion, claiming Anthropic could take “the next couple of GW from SpaceX” to support revenue growth; this capacity claim is commentary in the source rather than independently substantiated evidence.
- OpenAI said it is ending its partnership with Cursor following Cursor’s acquisition by SpaceX; under its proposal, Cursor’s direct access to OpenAI models would end on November 12.
- OpenAI said it will provide extensive transition support to developers affected by the decision.
- FactoryAI is offering all SpaceXAI and OpenAI models at 60% off for the next week; the post characterizes the two model families as complementary.
- OpenAI said it is ending its partnership with Cursor following Cursor’s acquisition by SpaceX, with a proposal to terminate Cursor’s direct access to OpenAI models on November 12; OpenAI also said it would provide extensive support to affected developers during the transition.
- GLM-5.3 delivers higher reported accuracy than GLM-5.3 Flash (69.0% vs. 63.4%), but costs substantially more per task ($3.99 vs. $0.24). GLM-5.3 also uses more tokens per task (80k vs. 73k), takes slightly more steps on average (125 vs. 123), and has slower time per step (17 vs. 12 seconds).
- For most workloads, start with GLM-5.3 Flash and use GLM-5.3 when the additional accuracy is worth the higher cost.
An AI application that depends on a third-party model provider may face strategic continuity risk if ownership or acquisition plans create conflicts: the post contrasts Cursor being cut off by OpenAI after an Elon acquisition with Windsurf being cut off by Anthropic after OpenAI tried to acquire it, concluding, “Not your weights, not your product.”
- OpenAI said it is ending its partnership with Cursor following Cursor’s acquisition by SpaceX; under its proposal, Cursor’s direct access to OpenAI models would end on November 12. OpenAI said it would provide extensive transition support to affected developers.
- Commentary argues that model labs will build model-specific harnesses and ecosystems while restricting access by other labs, making a lab-independent harness necessary for cross-model functionality; it concludes, “Long live LangChain.”
Isaac 0.5 model weights are officially available.
- Moore Threads says its MTT S5000 cards and MUSA stack provide day-one support for Zhipu’s open-source GLM-5.3-Flash (320B-A18B), described as the first native multimodal model in the GLM-5 series; the post reports an AA-index score of 57, matching Claude Opus 4.8. The development is framed as evidence that China’s domestic hardware-software stack can support a frontier-adjacent model at launch.
Terminal-Bench 4.0 updates the benchmark by calibrating per-task time, CPU, and memory resources, fixing tasks, and removing saturated tasks. Andy Konwinski says the project will treat benchmarks as software, with frequent version bumps for bug fixes, new tasks, and retiring saturated tasks; Harbor Hub provides tools to inspect what changed and why.
OpenAI said it is ending its partnership with Cursor following Cursor’s acquisition by SpaceX; under OpenAI’s proposal, Cursor’s direct access to OpenAI models would end on November 12. OpenAI said developers relying on its models in Cursor would be most affected and that it would provide extensive transition support.
- OpenAI said it is ending its partnership with Cursor following Cursor’s acquisition by SpaceX; under the proposal, Cursor’s direct access to OpenAI models would end on November 12. OpenAI said developers relying on its models in Cursor would be most affected and that it would provide extensive transition support.
Cursor has partnered with Anthropic since Claude Sonnet 3.5, and the relationship is expected to expand through increased compute support for Claude models in Cursor; the post also signals further Anthropic-related activity at SpaceX.
- Box and Nous Research partnered to integrate Box as a governed enterprise content layer for Hermes Agent; the official Box skill is now included with every new Hermes installation.
- The skill enables Hermes to manage and search Box files, extract document data, generate grounded answers with Box AI and Box Hubs, process approved sharing batches, and build Box-backed applications and webhooks through the CLI, Box Request, or official SDK.
- The integration includes enterprise safeguards: verifying the active Box user, preserving permissions, requiring confirmation for permission changes or destructive actions, and reading back writes to verify results.
𝕏 post by @thsottiaux
We unfortunately have decided that we cannot continue providing access to our models through Cursor and are ending our partnership. It boils down to trust and we’ve asked that this takes effect on November 12 to give you some time to plan.
Many have used the GPT models through Cursor and here are options we know should work in the future:
- We will continue to allow using your own OpenAI API key and similarly will continue to provide access through our IDE extensions for Cursor.
- We will keep working with the broadest range of tools and harnesses, some of which are OSS, but also many many closed-source ones.
We are as committed as ever to continue supporting developers and the flourishing ecosystem of tools, harnesses and products. We will also continue to invest in our own open-source initiatives and believe in broad optionality for developers.
You can read more about our decision in the blog: https://openai.com/index/our-decision-on-cursor-following-its-acquisition-by-spacex/ (opens in new tab)
We’re ending our partnership with Cursor following its acquisition by SpaceX. Under our proposal, Cursor’s direct access to our models would end on November 12.
We know that the people most affected by this decision are the developers who rely on OpenAI models in Cursor. We care about their experience in this transition and we’re ready to go above and beyond to support them.
- OpenAI announced that it would end its partnership with Cursor and stop providing access to its models through Cursor, citing trust; it asked for the change to take effect on November 12. It will continue supporting users’ own OpenAI API keys, access through its IDE extensions for Cursor, a broad range of other tools and harnesses, and open-source initiatives.
- Jerry Liu frames the move as part of a broader ecosystem split: frontier labs may seek end-to-end ownership of applications, harnesses, and models, while other AI companies favor access to mixed proprietary and open-weight models to optimize task performance and margins and reduce dependence on any one provider.
- The provider announced that it will end its partnership with Cursor and discontinue direct access to its models on November 12, citing trust concerns. Users will still be able to use their own OpenAI API keys and access the provider’s models through its IDE extensions for Cursor.