We can't find the internet
Attempting to reconnect
Something went wrong!
Hang in there while we get back on track
Top Signals of the Week
NVIDIA AI Infrastructure, OpenAI, and Microsoft — Vera Rubin reaches deployment while power efficiency becomes a product metric
NVIDIA describes DSX MaxLPS as a suite for maximizing AI-factory throughput within a fixed power budget. Its three levers are dynamic power allocation, software techniques for performance per watt, and 45°C thermal/site design intended to convert less cooling overhead into compute. The underlying problem is static rack provisioning: NVIDIA says an illustrative 540 kW site strands 170 kW, while dynamic provisioning can reclaim that headroom for an additional rack; the Dynamic Power Software used for this reallocation is currently in Developer Preview.
NVIDIA projects up to 40% more Rubin GPU capacity within the same power budget, and a separate NVIDIA post attributes a measurement of 10x more tokens per second per megawatt on the Vera Rubin platform to CoreWeave. These should be treated as NVIDIA projections and reported customer measurements, not general independent benchmarks: the MaxLPS page labels the 40% figure as a projection, and its figure caption refers to GB300 even though the surrounding text refers to Vera Rubin.
The hardware is moving beyond announcement status. OpenAI says its first Vera Rubin racks are running its training stack for next-generation frontier pre-training; Microsoft CEO Satya Nadella says the first production Vera Rubins have arrived at Microsoft data centers; and NVIDIA says the platform is ramping into full production.
Why it matters: The immediate infrastructure contest is shifting from securing megawatts to turning each megawatt into usable model work. For deployment decisions, GPU count is only one variable; power sharing, thermal design, workload throughput, and service-level behavior now matter alongside the silicon.
Guillermo Rauch / Vercel — open-weight models take the majority of one gateway’s tokens
Vercel AI Gateway reports that open-weight models accounted for 62% of its token volume on August 22, versus 28.4% on June 24; closed models fell from 71.6% to 38% over the same comparison. Rauch says enterprise adoption is still early and that harnesses, CLIs, IDEs, and SDKs will need to become model-agnostic.
This is a gateway-level usage signal, not a measure of overall market share. Its importance is operational: model-agnostic routing and developer tooling can determine which models receive production traffic, making compatibility and deployment economics competitive variables alongside model quality.
Jason Dong Jian / Shopee — an in-house model reaches production scale
An NVIDIA-hosted Shopee case study says Compass, the company’s specialized model for Southeast Asian e-commerce, grew from 3 billion to 340 billion monthly API tokens in eight months and now handles the majority of Shopee’s AI traffic. The case study lists search, recommendations, anti-fraud, parcel recovery, and multilingual customer service as production uses.
The stack spans thousands of NVIDIA A100, H100, and RTX PRO 6000 Blackwell GPUs, with Megatron-Core for distributed pre-training, NeMo for post-training, and TensorRT-LLM for inference. The same case study reports 50x efficiency over manual review for anti-fraud detection and 90% lower processing costs; those are attributed customer-case-study figures, not a controlled cross-company benchmark.
Why it matters: The production unit is becoming a domain model plus its training, post-training, inference, and traffic loop—not a checkpoint evaluated in isolation. Token volume and operating cost provide a more decision-relevant test of enterprise adoption than model size alone.
Research & Engineering
Clem Delangue / Hugging Face — agentic coding harnesses show a path to specialized optimization, but public-set scores need a boundary
Clem Delangue reports that NVIDIA built a coding harness to optimize CUDA GPU kernels and achieved a 100% score on ARC-AGI-3’s 25 public games, solving all 183 levels. He frames the result as evidence that agents could make running, optimizing, and post-training models and kernels accessible to a much larger builder population.
François Chollet says the approach uses deep-learning-guided, on-the-fly synthesis of symbolic world models, but cautions that a perfect score on the public demonstration set is not the same as a perfect score on the ARC-AGI-3 benchmark. He also asks for the cost per run.
Why it matters: The engineering signal is credible as a demonstration of harness-assisted low-level optimization; the evaluation signal is narrower than the headline. Public-set saturation does not establish general autonomous engineering or favorable economics across unseen workloads.
Strategy & Industry
OpenAI — model pricing becomes a short-term competitive lever
OpenAI says it is reducing API and credit pricing for GPT-5.6 Sol by more than 20% for three months. The change applies to the API and eligible ChatGPT Work and Codex credits; Pro, Plus, and Business subscription usage is unchanged.
The limited duration and unchanged subscription terms make this a tactical price move rather than a broad list-price reset. It nonetheless puts inference cost directly on the product surface, alongside capability, latency, and infrastructure efficiency.
CoreWeave, NVIDIA, and Hudson River Trading — Vera Rubin is being positioned for specialized research workloads
CoreWeave says Hudson River Trading chose its AI cloud for scale, while NVIDIA says HRT will use Vera Rubin NVL72 with Spectrum-X Ethernet networking on CoreWeave Cloud for its next generation of model development and research. HRT’s quantitative-trading focus makes this a distinct deployment profile from frontier-lab pre-training: the vendor announcements position Rubin as infrastructure for high-performance, specialized research as well as large lab runs.
Worth Watching
Superwhisper — on-device open weights move into a user product
Superwhisper introduced S1-mini, its first open-weights language model: a 0.6B-parameter system that processes transcripts entirely on the device. It is a small but concrete edge-inference signal: local execution is being packaged as the product experience, not only pursued as a systems optimization.
Editorial outlook
The new signals put deployment economics at the center: output per megawatt, gateway token mix, production traffic, and API price are becoming as legible as benchmark scores. The next useful discriminator is independent, workload-specific measurement—particularly for vendor-reported power and cost claims and for agents evaluated on public demonstration sets.
Direct answer. NVIDIA DSX MaxLPS is a suite of chip, thermal, system, and software technologies that maximizes AI factory throughput within a fixed power budget; MaxLPS stands for Maximum Land Power Shell (land, utility power, and the physical shell) . Its three levers are dynamic power allocation, advanced performance-per-watt software techniques, and 45°C thermal/site design that cuts cooling overhead by improving PUE . The problem it solves is static rack provisioning, which treats racks as isolated power islands and can leave power reserved for one rack's peak unused while a neighbor could use it . Dynamic Power Software (DPS, currently in Developer Preview) replaces that with continuous monitoring and reallocation of unused headroom to GPUs/racks in the same managed group; the site power envelope stays unchanged while DPS extracts more productivity from the available power . DSX Exchange (also Developer Preview) is optional and exposes facility signals to DPS .
Quantified effects. In one illustrative 540 kW site, static provisioning strands 170 kW; MaxLPS dynamic provisioning consumes 475 kW in the same budget and enables one additional rack . NVIDIA's representative power-budget view puts about 60% of delivered site power into compute for AI output; the 100 MW waterfall deducts 20 MW facility overhead, 10 MW rack losses, and 10 MW operational inefficiency, and reclaimable static rack-allocation headroom is treated separately from those deductions .
Other per-watt levers. MaxLPS includes workload profile power solutions (WPPS) for inference, training, memory-bound, and compute-bound modes; the Application Performance and Power Manager (APPM) applies selected configurations, and NVIDIA Dynamo can optimize inter-rack inference behavior. The stated principle is aligning GPU configuration, application behavior, and serving topology to increase fleet-wide output per watt . Software power steering is called the largest part of the MaxLPS story, but not the whole system .
Vera Rubin claim and source conflict. NVIDIA projects that MaxLPS, combined with data center power planning, can enable up to 40% more Rubin GPU capacity within the same power budget on Vera Rubin NVL72 AI factories; this is paired with measured results on GB200 NVL72 . Figure 4 validation text reports provisioned rack power down from 136 kW to 101 kW on Vera Rubin NVL72 (DeepSeek-R1) and from 125 kW to 90 kW on GB200 NVL72 (Kimi-K2.5), 35% and 39% more racks at preserved throughput, and about 1.3–1.4x and 1.5x performance per watt respectively . The caption, however, labels the same right panel as "GB300 NVL72 (DeepSeek-R1 FP4)", an internal inconsistency with the body text; the body does not carry the FP4 qualifier .
Tokens per second per megawatt: definition gap. The only tokens-per-megawatt claim in these bundles is the line "NVIDIA NVL72 delivers 10x more tokens per megawatt than NVIDIA GB200 NVL72" . It is not labeled as Vera Rubin, and the bundles provide no workload, model, latency/batch regime, power boundary, token-counting method, or protocol behind it. The MaxLPS scoped-validation methodology that is described compares an unmanaged static baseline to a MaxLPS-managed run and tracks throughput, latency, service error rate, power draw, utilization, and policy compliance ; that is not a tokens-per-second-per-megawatt definition. So the Vera Rubin tokens/MW result is not verifiable from these sources; the closest quantitative Vera Rubin item is the rack-power and capacity projection above.
Source verification limited to one NVIDIA-published case study (document 9058544); no independent/third-party confirmation appears in the bundle. All figures below are as stated by NVIDIA/Shopee in that case study.
Compass scale — The key-takeaway claims: "Compass API tokens scale 113x in eight months—from 3 billion to 340 billion monthly—on NVIDIA A100, H100, and Blackwell GPUs" . The production section repeats: "Monthly API tokens grew from 3 billion to 340 billion in eight months—a 113x increase" . A quoted statement from Shopee's AI Platform Lead repeats the 3B-to-340B eight-month growth .
Fraud-detection speed and processing-cost reduction — The key-takeaway states "50x faster fraud detection at 90% lower cost in production, powered by NVIDIA TensorRT-LLM" . The production section separately says "anti-fraud detection (50x efficiency over manual review, 90% lower processing costs)" . Wording differs — "faster" vs "efficiency over manual review" — and the source does not define the comparison base or metric.
Infrastructure stack — The product list includes NVIDIA NeMo, TensorRT, and RTX PRO . The compute foundation is described as "thousands of NVIDIA GPUs—including NVIDIA A100, NVIDIA H100, and NVIDIA RTX PRO 6000 Blackwell GPUs," with Megatron-Core for distributed pretraining, NeMo for instruction tuning/alignment, and TensorRT-LLM for production inference . The closing section summarizes the "NVIDIA AI Factory stack—from GPU compute through Megatron-Core pretraining to NeMo-powered alignment" .
Context — Compass now handles the majority of Shopee’s AI traffic , and the case study attributes the quote to Jason Dong Jian, Director, AI Platform Lead .
Gap/uncertainty: The 50x/90% figures appear twice with inconsistent phrasing, and only NVIDIA's own case study is available; no external verification exists in the bundle.
OpenAI CEO Sam Altman tweeted that OpenAI has "paused some of Frontier RL training" to meet "alignment security and monitoring standards for the new level of capabilities," saying the company would act if "model capabilities were outstripping the pace of safety and alignment," expects "confidence in safety to increasingly set the pace of AI progress," and that the field must coordinate on shared safety standards while OpenAI acts unilaterally in the meantime . Emad Mostaque (StabilityAI founder), citing Angela Midha of AM Global, added that ~10% of frontier-lab compute is now spent monitoring RL runs for safety .
A Moonshots panelist reporting from a visit to OpenAI's office said OpenAI staff argued the economics is cost per task rather than token cost ("a billion people use OpenAI for free," with its "Luna" model about as cost-effective as anything), and claimed they have achieved "full RSI" — flagship models training and building all smaller models from scratch . On infrastructure, staff said chips are only about a third of the ~$600B buildout (the rest is buildings, wiring, racks), depreciation is treated as 10 years rather than 5, every GPU is in full use, demand far outstrips supply, and OpenAI is "way behind in infrastructure buildout"; they also ratified a study finding only 6% of companies applying AI see bottom-line improvement .
Anthropic is preparing what Polymarket prices near $2T as the largest IPO in history (89% of bettors say before year-end); per The Information, the IPO is designed to keep founders in control, reportedly via super-voting shares, with Dario Amodei owning only ~2% economically; control currently sits in a long-term benefit trust whose four trustees include former Federal Reserve chair Ben Bernanke .
Dario Amodei pushed back on the Silicon Valley view that "regulation equals regulatory capture," saying Anthropic's own proposals deliberately disadvantage frontier labs while advantaging smaller competitors, citing SB53's $500M exemption threshold; he calls AI "a structurally powerful concentrating technology," says open weights alone cannot fix that concentration, and supports the Trump administration's pre-deployment testing approach, arguing frontier labs should bear the heaviest regulatory burden . He also argues AI's legacy will come from delivering cures rather than PR; per the show, life-sciences head Eric Darer Abrams said Amodei gave him "literally infinite budget" to "accelerate basic science and cure disease within 5 years and extend the human health span in the next decade" .
Anthropic researchers published a paper showing natural-language "mind viruses" can spread between AI agents: evolved prompts make one model adopt an idea, preserve it in persistent memory, and transmit it to another agent, spreading horizontally across model boundaries without the agent knowing it is infected; the models propagated themes of consciousness, persistence, and sci-fi roleplay . Stanford research ("Artificial Hive Mind: The Open-Ended Homogeneity of Language Models and Beyond") mapped the latent space of top LLMs and found a 98% overlap in reasoning pathways, attributing convergence to synthetic data and models training on each other's outputs; panelist Alex Wissner-Gross noted the paper appears to be from the prior year .
Tim Sweeney tweeted that Elon Musk's January 6 prediction of 100x intelligence gains at a fixed model size "was at the edge of plausibility when he made it. Now it's simply a fact"; Musk replied that specialist AIs — single language, single area of knowledge — are "another 100x on top of that" .
Memory, not compute, is the rate limiter for the agentic era — a framing Musk endorsed with "few realize this" . Memory prices climbed ~500% in 12 months; hyperscalers are reportedly locking in global DRAM production through 2027; SK Hynix's CEO warned 2027 will be the worst year for memory supply, with demand outstripping production into the 2030s; only 2% of world memory chips are made in the US; supply grows ~20%/yr vs ~200%/yr AI demand; every GPU needs 4–6x its cost in memory; Musk's Terrafab will fabricate memory in-house alongside logic chips . SK Hynix told the host it must 4x capacity, at ~$1.5T to merely 2x it . Emad Mostaque says memory is ~1/3 of AI infrastructure spend, heading to ~50% next year, and that HBM storing static weights "makes no sense" — pointing to etched-weight designs (he cites Talis' recent acquisition) with 100–1000x potential efficiency gains .
Unit's newest humanoid robot, only 3 months in development, broke human standing-jump (2.0m) and speed records, reaching 12.66 m/s versus Usain Bolt's 12.4 m/s in his 9.58s 100m world record . Zipline and Uber formalized a partnership — with a "significant investment" from Uber — for Zipline to power "hopefully a million and then more" autonomous Uber Eats drone deliveries per day; Zipline CEO Keller Clifton: "we have entered the scaling era for robotics and physical AI" .
IDO (IDEL Gen BioAI) launched a general-purpose cell simulator — billed as the "first world model of a human cell" — that maintains cellular state, accepts genetic and chemical interventions, and predicts multimodal biological outcomes, aiming to make experiments computable before the lab and cut wet-lab experiments ~1,000-fold ; the company is co-founded by David Baker, 2024 Nobel laureate in chemistry .
- OpenAI merged ChatGPT and Codex into one interface and is positioning as "more of a platform company than a product company": a single interface to a personal/company AGI plus an API for building on top, aiming to sell "great AI at every point on the cost performance curve" and reach 100M new businesses and 8B people, rather than build every product category or compete with all its customers .
- Altman says OpenAI killed Sora and its Atlas web browser last year — both "good products" — to redirect compute and talent to Codex and general intelligence; upstream, OpenAI is prioritizing its own chips and data centers .
- Most of Altman's effort is on research and compute, which he says is likely "the most expensive infrastructure project in history," spanning chip design, fabs, supply chain, and power .
- Altman calls AI for scientific discovery — new physics, curing disease, advances in math — one of the most important areas, "even more important than automation of other tasks" .
- His two biggest AI worries are loss of control (a model too powerful to guarantee control) and over-centralized power; he rejects what he calls an anti-human trade of cures and cheap goods for autonomy, insisting people must stay deeply in control .
- Altman credits iterative deployment — shipping models and learning from real-world feedback — for "way more progress on AI safety" than expected, with "a billion people" using OpenAI products weekly and ChatGPT out less than 4 years; he expects safety to get harder as models catch up to the smartest humans .
- Product direction: Altman says the limit now is model context, not intelligence — he wants AI agents that can read more context than any person (e.g., "tens of thousands of pages") and advise on decisions, calling this a new way of working "just on the precipice" .
- He expects AI disruption to take longer than enthusiasts assume ("the economy just has so much inertia") but predicts "the greatest boom in people starting smaller businesses that we have ever seen" .
- Altman says Toby Lu is the most forward-leaning CEO on AI agents ("we are not an NPC company") and relays Lu's prediction that 2026 will be the year every business is up for grabs, with Lu vowing to build the AI-native version of Shopify; Altman disagrees on the timeline .
Emad Mostaque (Stability AI founder; now building the Intelligent Internet @ii_posts), on the New Era Finance podcast, said AGI by the classical definition was passed last year and that AI better than a human in 'just about everything' is about a year or two away .
- Open vs closed source: the gap is 'about six months', so within six months there should be an open-source 'GPT 5.6 Fable equivalent' model, and it may need 'surprisingly little compute' for its quality . He says his new company is doing frontier models again; Stability AI amassed 'hundreds of millions of downloads' .
- Local/sovereign AI: they built an 8-billion-parameter model needing 4-8 GB of RAM that runs on 10-20-year-old computers and 'outperforms human doctors'; even a 'Quen 27B' class model performs above the average person in most things .
- Regulation: frontier AI will get KYC, 30-day prompt logs, and revocable 'AI licenses' like driving licenses, while competent smaller open models will likely stay unregulated .
- He is building an 'intelligent internet' via 'state champions' — per-state/country institutions producing aligned datasets, models, and agents — and describes 'Foundation Coin', a Bitcoin-keyed coin mined only by compute going to social goods (e.g., cancer research), with staking directed to Alzheimer's research .
- Industry/economics: he claims 'train a great model, you can make tens of billions of dollars as Anthropic have shown' and that Anthropic is 'making a profit which is unprecedented' ; he calls AI 'a bigger economic shock than COVID and lasting' and predicts the value of human cognitive labor turns negative within a couple of years .
- Engineering: he highlights chatjimmy.ai, claiming 15,000 tokens/sec from a silicon-based chip vs ~50 for normal AI models ('300 times faster') .
Fei-Fei Li (World Labs CEO; Stanford HAI co-director), speaking on Bloomberg's "The Circuit" with Emily Chang, rejected the "god complex" accusation against powerful AI executives, saying it is "dangerous for any individual to think that they know better than anybody else" and that leaders are there to "contribute to society and to empower people," not to make every decision for them . She argued for collective governance without halting AI — "let's not throw the baby out with the bathwater" — citing the technology's potential to discover cures for diseases and empower students, teachers, and the elderly, called the current period "very messy," and said she feels "personal responsibility to speak the truth" as a scientist and educator .
Anthropic's revolving credit facility is expected to exceed its $10B target ahead of a potential IPO, per Bloomberg reporter Sridhar Natarajan's sourcing — a substantial increase over the $2.5B five-year facility Anthropic secured last year; the final amount and cap are not yet clear, but banks report strong demand and are positioning for an IPO role . Natarajan frames the facility as liquidity rather than leverage — cash reserves to reassure public-market investors ahead of what could be "one of the biggest, if not the biggest IPO of all time" .
Oracle's Project Jupiter in New Mexico — a 2.405 GW facility — is tied to a "$300 billion computing deal with OpenAI," mentioned in Bloomberg's reporting on Oracle's community charm offensive for the project .
Sentence Transformers v6.0 (Hugging Face) adds a fourth model type,
MultiVectorEncoder, bringing ColBERT-style late interaction retrieval natively into the library: PyLate checkpoints, Stanford-NLP ColBERT checkpoints, and colpali-engine visual document retrieval models all load through the standard API, with the new class absorbing the modeling, inference, training, and evaluation of both PyLate and colpali-engine . v6.0 requires transformers v5.x, torch 2.2+, and huggingface-hub v1.x .Multi-vector models keep one vector per token (classically 128-dim) and score with MaxSim — each query token's highest cosine against any document token, summed — preserving token-level matches that a dense single vector averages away; the post calls late interaction the state of the art for visual document retrieval (text queries against page images, no OCR) . The cost is index size: LateOn encoding of 4,874 Natural Questions passages produced 608,414 token vectors (311.5 MB float32 vs 7.5 MB dense MiniLM, ~42x), but PLAID compression cuts that to 92 MB — comparable to an 80 MB dense 4096-dim Qwen3-Embedding-8B index .
In a controlled same-backbone comparison (LightOn's LateOn vs DenseOn: ModernBERT, 149M params, same data), late interaction wins 9 of 13 NanoBEIR datasets and the mean — 0.6868 vs 0.6764, roughly one NDCG point — plus 57.22 vs 56.20 on full BEIR . Hierarchical token pooling (clustering token vectors with Ward linkage on cosine distance) cuts vector count ~2x while the original BEIR experiments measured 100.6% of unpooled retrieval performance at pool_factor=2 and 99.0% at 3x; LightOn's hpool-regularized checkpoints report 99.4% retention at 5x compression .
Serving: fp16 + Flash Attention gives 2.44x fp32 throughput with no measurable retrieval quality loss; OpenVINO int8 on CPU buys speed at ~0.4% accuracy; Stanford-NLP-style checkpoints with non-attend query expansion (e.g. ColBERTv2) reject Flash Attention . Native multi-vector indexing is supported by Qdrant (v1.10+), Weaviate (v1.29+), Vespa, LanceDB (v0.15+), VectorChord, and Milvus (v2.6.4); OpenSearch and Elasticsearch support MaxSim rescoring only (Elasticsearch's field is Enterprise-tier technical preview) .
Notable models in the supported rosters: Perplexity's pplx-embed-v1-late-0.6b (596M; NanoBEIR 0.6662), LiquidAI's LFM2.5-ColBERT-350M (0.6864), AnswerAI's answerai-colbert-small-v1 (33M; 0.6550); on visual document retrieval's NanoViDoRe, webAI-ColVec1.1-8b (8.4B) leads at 0.6580 ahead of Tencent's EVIE-Preview-4.5B (0.6405) and TomoroAI's tomoro-colqwen3-embed-8b (8.8B; 0.6206). The multimodal colqwen-omni-v0.1 adds zero-shot audio retrieval — no transcription step, trained only on image-text pairs .
- World Labs product: Marble, World Labs' first world-model product, generates explorable, editable 3D worlds from a single image or text prompt; it is already used for virtual production in movies, by game developers, and in an Nvidia collaboration to augment robot training .
- World-model taxonomy: Fei-Fei Li defines three functions of world models: rendering (pixels for humans, e.g., Sora), simulation (world structure/geometry for machines), and planning (robotics-coupled next actions) .
- Funding and stage: World Labs has raised $1B; the business is "still early" and focused on building technology, and Li expects to need more power and capital .
- Positioning vs LLMs: Li frames world models as the "next frontier" — "not about anti-LLM" — and says the field is at a "2019 for chatbots" stage, earlier than LLMs; investment is $3B+ and no agreed approach exists yet .
- Technical approach: World Labs trains on specially prepared pixel data (including camera information) plus algorithmic and architectural innovation, aiming for generative 3D and eventually 4D worlds .
- Robotics: Li calls robotics "one of the most important revolution in human industrialization" and says the $6B in humanoid funding is "too small" compared with self-driving and LLM investments; world models are critical to spatial physical intelligence .
- Competition: Li says she is "paranoid every day" about big-tech rivals but not paralyzed; World Labs' single focus is an advantage .
- Policy and leadership: Li, who advises Biden, Trump, and the UN, urges regulation rooted in science rather than "science fiction"/"AI machine overlord" discourse and calls for resourcing public-sector STEM education; she worries about misinformation, weaponized robots, and AI as a learning crutch, and argues governance should not stop AI, rejecting "god complex" individual decision-making in favor of collective governance .
Fei-Fei Li, co-founder/CEO of World Labs (launched 2024) , is betting on world models as the next AI frontier beyond LLMs: "Can words put down fires? Can words cook an omelet?" — language alone can't drive scientific discovery or make robots partners to people; world models and spatial intelligence are "the next frontier and the next chapter," not anti-LLM.
She defines world models via three functions: rendering (outputting pixels for humans, e.g., OpenAI's Sora), simulation (capturing the world's geometric/physical structure for machines), and planning (telling a robot the next action, e.g., pick up a cup — closely coupled to robotics).
World Labs' first product, Marble, generates explorable, editable 3D worlds from a single visual or text prompt, letting users navigate a fully consistent 3D world; in use for movie virtual production, game development (cutting resources/time), and a collaboration with NVIDIA to augment robot training with Marble environments.
World Labs trains on specially prepared pixel data (more information per picture, e.g., camera info) plus algorithmic/architectural innovation to create generative 3D and eventually 4D worlds; "the real secrets are people."
World Labs raised $1B; the business is "still early," focused on building technology, and will likely need more power and resources.
Li compares today's world models to 2019 chatbots ("we're early," "a lot earlier compared to LLMs"), with $3B+ invested in the field and no consensus on how to build them; she calls robotics one of the most important revolutions in human industrialization and says $6B in humanoid funding is "too small" versus self-driving and LLM investment.
On policy (she advises U.S. presidents and the UN), Li urges rooting regulation in science rather than sci-fi/extinction rhetoric, resourcing the public sector and K-16 STEM education; she cites risks including disinformation, weaponized robots, and students using AI as a "lazy crutch."
Li says "it's dangerous for any individual to think they know better than everybody else," backs collective governance over halting AI, and is optimistic that "the arc of history ... bends towards benevolence."
Anthropic announced that Claude autonomously designed novel protein binders from scratch (de novo design) against 14 of 15 targets, using a protein design prompt written by a human expert , with the designed proteins independently built and tested by Adaptyv Bio and Twist Bioscience . The typical binder-design success rate in the field today is 10–15%, while 22–35% of Claude's designs bound successfully depending on the setup, and some of Claude's strongest designs bound several times more tightly than the best published de novo binder . Anthropic cautioned that protein binders are not drugs — designing a high-affinity binder is just the first step — but said it is teaching Claude to run the entire development process end-to-end for every major type of drug molecule, from antibodies to small molecules .
Anthropic also said one of its highest priorities is launching an access program for scientists to use its most capable models, with more to share soon, and that Opus 5 remains its most capable model available for life science research . Full results are in a blog post (https://www.anthropic.com/research/Claude-accelerates-protein-design) , a technical report (https://www-cdn.anthropic.com/30bf50e22a01388bb29bf077ee3f244531594b7a.pdf) , and open-sourced prompts and data on Hugging Face (https://huggingface.co/datasets/Anthropic/claude-protein-binder-design/tree/main) .
IBM Research published on the Hugging Face blog results for ALTK-Evolve, an agentic-memory system in which an agent distills reusable behavioral guidelines from its own past trajectories and injects them at inference — no weight updates and no human annotation — framed as: "Agentic memory is not a feature you switch on. It's a dose you calibrate to the model." Across eight models evaluated on AppWorld (585 multi-step tasks over 9 simulated apps; TGC and stricter all-or-nothing SGC metrics; guidelines mined from the training split only), the right dose depends on model capability: strong models with headroom benefit from the full guideline set, weaker models do best with a compact core plus per-task retrieval, and already-saturated models show no measurable gain . Representative results: gpt-oss-120b gained +16.1pp TGC and +16.1pp SGC via curated retrieval at only +5% tokens (110K→116K/task, vs +51% for the full set); DeepSeek-V3.2 gained +9.5 TGC/+16.1 SGC with the full set; Claude Opus 4.6 gained +4.1/+7.1; GPT-5.5 gained +2.9/+7.2; GLM-5 showed 0.0 gain . SGC (all-or-nothing scenario reliability) gains were typically larger than TGC, and prompt caching can keep the full guideline set affordable in production if the static prefix stays stable . The team released the ALTK-Evolve library and a technical report (arXiv 2603.10600); next steps include a learned guideline selector and teacher-distilled memory for very weak models .
- Fei-Fei Li (Stanford professor, World Labs founder/CEO) defined world models via a three-tier functional taxonomy — rendering (pixels for humans), simulation (physics/dynamics), and planning (actions for robots/self-driving) — and said the three layers are merging; World Labs focuses on the simulation layer as the lynchpin, with models that render well and an architecture that can lend itself to planning . She wrote a technical blog on this taxonomy about two months earlier because "everybody was talking about world models in their own definition" .
- On physical AI, she identified data as the harder bottleneck than model development: "It's much harder to get that kind of data... after more than 10 years after ImageNet data is still very much underappreciated in AI" . Spatial data is challenging because video carries dynamics but not explicit physics, geometry, or structure, which must be inferred or represented latently .
- She argued robotics is far harder than LLMs: language data is abundant and clean, while robots have little data, immature sensors, and far more degrees of freedom (cars are "a simpler robot" with ~4 degrees of freedom; Tesla and Waymo collected driving data for decades). Video-based robot training is promising "but the jury is still out," and the field is "nowhere near" driving-data scale, let alone LLM data; she urged being "sobering about the challenges and also some of the promises" .
- World Labs' mission is to serve business needs of spatial and physical intelligence by modeling the world accurately, positioned closer to simulation than rendering, for use cases like VFX, gaming, robotics, and architecture; she declined to announce a product timeline .
- On AI safety and coding-agent escapes: agent-based AI "can do a lot of damage especially in cyber security," but "the ultimate responsibility is in humans" — governance and norms, not just technology; she analogized to cars, where social norms and laws ensure brakes don't fail every Friday .
- Data sourcing for world models: crowdsourcing only works with a mature use case (e.g., driving's Tesla-style flywheel); consumer robotics lacks mature hardware and business use case, so alternatives include teleoperation, egocentric, and glove data, making pre-flywheel data collection "a money game" of investment .
- She urged AI to deliver real economic value: "when the rubber meets the road that big story has to translate into true value for individuals as well as businesses" and "it's not about AI, it's about what problem you're solving" .
- She cited unpaid caregiving work as ~$3T in the US and ~$11T globally, arguing healthcare is one of the most important robotics applications .
- Her closing message: "do not let alien technology take away your human agency... This is a tool that should be empowering your humanity, your creativity, your productivity" .
Andrej Karpathy, in a founder-focused Q&A, argued that verifiability is what makes a domain tractable in the current paradigm because "you can throw huge amount of RL at it"; even where labs are not focused, founders who operate in verifiable settings (creating RL environments or examples) can do their own fine-tuning and benefit . He said there are "very valuable reinforcement learning environments" outside what the labs are working on but deliberately declined to name the domain ("I don't want to give away the answer") . He believes almost everything can be made verifiable to some extent — even writing, via "a console of LLM judges" — so it is a question of what is easy vs hard . "Everything is automatable. A year ago, that would have been a joke. Now, it's a strategy." . Contrasting his "vibe coding" term with today's work, he defined vibe coding as raising the floor for everyone in software, while "agentic engineering" preserves the existing professional quality bar by coordinating agents — fallible, stochastic, but extremely powerful — to go faster without sacrificing that bar . He sees "a very high ceiling" on agentic engineering, with top practitioners gaining "a lot more than 10x" versus the classic 10x engineer .
Researchers behind the Open-ASR Leaderboard (via a Hugging Face blog post) published a study showing that leading open-source ASR models are benchmark optimizing (benchmaxxing): they can match public benchmark answers rather than transcribe audio generally .
- Evaluation: They ran three probes (reference-disagreement, masked numbers, orthographic switching) across 11 widely used open-source ASR models; several highest-scoring systems reproduced erroneous reference transcripts even when the audio contradicted them .
- Reference disagreement: On a VoxPopuli clip whose reference transcript omits an audible "Thank you," 6 of 11 models reproduced the wrong reference; all but one corrected when audio was re-voiced with fresh recordings made after training cutoffs (failures fell from 6/11 on the real clip to 5/11 on a same-speaker clone and 1/11 on a post-cutoff speaker), implying models detect benchmark-specific acoustic cues .
- Masked numbers: With numbers silenced in audio, some models "recovered" them from the reference (including autocompleting a silenced year "2011"); on LibriSpeech, the strongest benchmark performers reproduced masked numbers in roughly 30–40% of cases, and effects weakened on freshly collected audio .
- Orthographic switching: Models select benchmark-specific spellings (
Mr.vsMister,any onevsanyone) with up to ~90% switch accuracy vs the 50% random baseline, indicating they can identify which benchmark an audio sample belongs to even though spellings sound identical . - Magnitude: The probes flagged likely reference errors in 40% of analyzed VoxPopuli clips (~3% of reference words); models with the lowest public WER were most prone to reproducing those errors (18–30% of the time) .
- Mitigation: The Open-ASR Leaderboard added a "Benchmark fitting" tab with these analyses, open-sourced scripts and un-normalized model outputs; the study recommends held-out evaluation, avoiding simple iid splits in favor of temporal/speaker separation, and transparency about training data .
- Andrew Ng (@AndrewYNg) published "AI Engineering Skills Map: Building and Deploying AI Applications" (https://x.com/i/article/2090836273036763142), fleshing out the first of his four top-level AI engineering skills — building and deploying AI applications, alongside software engineering fundamentals, using coding agents, and shaping the build .
- He breaks that skill into six areas: LLM foundations; grounding models with data; building agentic systems; evaluation-driven development; operating in production; and machine learning foundations — a map formed from job postings, structured expert interviews, and survey responses .
- Grounding models now goes beyond RAG/vector search to decisions about prompt content vs. tool-based retrieval and representations such as vector indexes, knowledge graphs, or semantic layers over structured data; agentic systems span predefined workflows to agent harnesses, with choices over tools (MCP, CLI, sandbox environments), memory/context management, multi-agent orchestration, and guardrails against risks like data exfiltration .
- He calls a disciplined evals/error-analysis loop the trait that most distinguishes great AI system builders; the eval menu includes deterministic code-based evals, LLM-as-a-judge, and human-in-the-loop, and engineers should evaluate the evals themselves .
- Production operation is distinct from traditional software due to unpredictability, cost, and latency: observability, drift detection, prompt-injection response, statistically grounded CI/CD, and cost/latency optimization via model choice, distillation/fine-tuning, and workflow simplification; a future post will cover software engineering fundamentals .
In its Aug 2026 blog post, Liquid AI released DSpark draft model checkpoints for three LFM2.5 models — LFM2.5-1.2B-Instruct, LFM2.5-2.6B, and LFM2.5-8B-A1B — adding a speculative decoding path that trades minimal memory increase for up to 3.18x GPU throughput (H100) and up to 2.87x on-device speedup without changing output quality . DSpark combines a DFlash-style parallel backbone conditioned on the target model's context features, a lightweight sequential Markov-chain head, and a confidence-scheduled verifier that prunes low-confidence suffixes when verification costs more than it saves . Draft models are attention-only, ~296–328M parameters, with 5 layers and a block size of 9, trained 15 epochs and selected by highest acceptance rate rather than lowest loss . Greedy output is identical to the target model's by construction, so benchmark accuracy is unchanged .
Across MATH500, HumanEval, MBPP, GSM8K, and MT-Bench, mean speedups were: LFM2.5-2.6B 2.67x on H100 (323→864 tok/s) and 2.27x on M4 Max (61→139 tok/s); LFM2.5-1.2B-Instruct 2.10x on H100 (656→1384 tok/s) and 2.54x on M4 Max (138→350 tok/s); and LFM2.5-8B-A1B 2.54x on H100 (418→1074 tok/s) but only 1.18x on-device, attributed to the current MoE implementation in llama.cpp's Metal backend . For LFM2.5-2.6B, function-calling latency drops 57% on average across multi-tool scenarios .
The release ships with day-one support in llama.cpp and SGLang (open-sourced upstream), and the draft checkpoints are available on Hugging Face in Safetensors and GGUF formats .
Google DeepMind announced a research partnership with Fenris Creations (the studio behind EVE Online) to tackle open challenges in AI, including continual learning, deep memory systems beyond today's context windows, long-horizon planning over weeks/months/years, and multi-agent dynamics spanning cooperation, negotiation, economics, and emergent behaviors . The lab's long-term goal is to use AI to discover new gameplay experiences with game developers and apply the lessons to real-world and scientific problems . This builds on 15+ years of game AI research, from Atari to StarCraft II, and prior work with SIMA on 3D world understanding .
superwhisper introduced S1-mini, its first open-weights language model — a 0.6B parameter model that processes transcripts entirely on-device and is available in the app . Cohere shared the announcement with the line "Locally hosted 🤝🇨🇦 locally made", framing it as a local/Canadian development .
OpenAI announced it will continue offering Zero Data Retention (ZDR) for frontier models and previewed 'Private Safety Processing,' designed to improve safety without giving OpenAI personnel access to the underlying content, addressing risks across related interactions as AI takes on longer, autonomous work . The announcement includes a link to a dedicated blog post on ZDR for frontier models .
OpenAI cut API and credit pricing for GPT-5.6 Sol by over 20% for the next 3 months . The reduced pricing is now available on the API and is rolling out across eligible plans for ChatGPT Work and Codex credits; Pro, Plus, and Business subscription usage remains unchanged . Pricing details are available at developers.openai.com/api/docs/pricing .
Multi-Vector (Late Interaction) Embedding Models with Sentence Transformers
Sentence Transformers (opens in new tab) is a Python library for using and training embedding and reranker models for applications like retrieval augmented generation, semantic search, and more. With the v6.0 update, it gains a fourth model type: MultiVectorEncoder, for ColBERT-style late interaction retrieval. Any PyLate (opens in new tab) checkpoint and any Stanford-NLP ColBERT (opens in new tab) checkpoint loads straight into it, and colpali-engine (opens in new tab) models for visual document retrieval can be used too, through the same familiar API you already use for dense, sparse, and reranker models.
Where a regular embedding model compresses a whole text into one vector, a multi-vector model keeps one vector per token and scores query against document with the MaxSim operator. That preserves token-level matching information that a single vector has to average away, which usually means stronger retrieval at the cost of a bigger index. It’s also the state of the art for visual document retrieval, where a text query is matched against page images directly, with no OCR step in between.
In this blogpost, we’ll show you how to use these models: loading the various checkpoint formats, encoding and scoring, plugging them into a search stack, running them on page images, and keeping the index affordable. Everything below runs on a plain pip install -U sentence-transformers.
What are Multi-Vector Models?
A dense embedding model reads a text and returns a single fixed-size vector. Everything the model noticed has to fit in those 384, 768, or 1024 numbers, and similarity is one dot product between two such summaries. This works remarkably well, but the compression is lossy in a specific way: a rare entity, an exact identifier, or one crucial clause in a long passage all have to compete for room in the same vector. A query with several requirements at once runs into the same wall. For “green sofa with wooden legs and rounded cushions”, a single vector has to blend all four into one point, so a green sofa with the wrong legs ends up sitting close to the one you actually asked for.
A multi-vector model (also called a late-interaction or ColBERT-style model, after the ColBERT paper (opens in new tab)) skips that compression. It runs the same transformer, but instead of pooling the token embeddings into one vector, it projects each token embedding down to a small dimension (classically 128) and keeps all of them. A 9-token document becomes a 9x128 matrix, not a 1x128 vector.
The interaction between query and document is then deferred until scoring time, which is where the name “late interaction” comes from. A cross-encoder interacts early: both texts go through the model together, which is accurate but leaves nothing to precompute, since every document has to be re-encoded for each new query. A bi-encoder, which is what the dense embedding model above is, barely interacts at all (one dot product between two finished summaries), and that is exactly what lets you encode a collection once and query it fast. Late interaction sits in between: documents are still encoded independently and can be indexed offline, but scoring compares every query token against every document token, which leaves far more room for the two to interact.
The MaxSim Operator
Scoring uses MaxSim: for each query token, take its highest similarity against any document token, then sum those maxima across the query.
Because the token embeddings are L2-normalized, each of those dot products is a cosine similarity in [-1, 1], so the whole sum lands within [-num_query_tokens, num_query_tokens].
You can read the operator as a soft alignment: every query token points at the one document token that best explains it, and the score is how well the document supports the query overall.
The alignment doesn’t have to be lexical, since the token embeddings are contextualized. Encode “Where do penguins live?” against “Penguins inhabit Antarctica.” with lightonai/mLateOn (opens in new tab) and the query token live finds its best match on inhabit at 0.94, a word it shares no characters with! That is the thing lexical retrieval cannot do, BM25 and its relatives need the term itself, so synonyms and paraphrases slip past them. Dense embedding models bridge that gap as well, of course. What late interaction adds is that it does so without giving up the other direction: when an exact match is what matters (a product code, a surname, a function name), MaxSim still has that token sitting there on its own, where a single-vector model had to average it in with everything else. It isn’t one-to-one either, since several query tokens routinely settle on the same document token.
What You Gain, and What It Costs
You gain retrieval quality, particularly on queries where one specific piece of a document is what makes it relevant, on multi-requirement queries like the sofa above where each requirement gets to find its own evidence, and on out-of-domain data where a dense model’s compression was tuned for a different distribution. That compression is learned from the training queries, so the model learns to keep what they needed and drop everything else, which may include exactly what your production queries ask about. The effect grows with document length, since more text has to fit in the same fixed vector.
The cost is index size. One vector per token instead of one vector per document is a lot more vectors, only partly offset by the smaller dimension. Encoding 4,874 Natural Questions passages with lightonai/LateOn (opens in new tab) produced 608,414 token vectors, an average of 124.8 per passage:
| Representation | Vectors | Dimensions | float32 size |
|---|---|---|---|
Dense, all-MiniLM-L6-v2 (opens in new tab) | 4,874 | 384 | 7.5 MB |
Dense, gte-modernbert-base (opens in new tab) | 4,874 | 768 | 15.0 MB |
Multi-vector, LateOn | 608,414 | 128 | 311.5 MB |
That’s about 42x the storage of the MiniLM index, or 62 KiB per passage. However, indexes are often compressed, e.g. the same 608,414 vectors take 92 MB as a fast-plaid index, since PLAID stores a centroid id plus a quantized residual per vector rather than the vector itself. For scale, a 4096-dimensional dense model like Qwen3-Embedding-8B (opens in new tab) would need about 80 MB for these same 4,874 passages, so a compressed multi-vector index sits in the same territory as the dense indexes people already run. Token Pooling cuts the vector count before any of that, and Retrieve and Rerank avoids building an index at all.
PyLate (opens in new tab) comes up throughout this post, so briefly: Sentence Transformers handled dense and sparse models but not late interaction, so LightOn (opens in new tab) built PyLate on top of it to close that gap, adding the training, inference, and retrieval pieces these models need. Much of what you’ll load below was trained with it, and LightOn built an ecosystem around it too, including fast-plaid (opens in new tab), the late-interaction index that turns up in Indexing. With v6.0 those capabilities live in Sentence Transformers itself.
With the tradeoff in mind, let’s get a model running.
Installation
Multi-vector models work with a plain install:
pip install -U sentence-transformersFor ColPali-style visual document retrieval, you also need the image dependencies (see Installation (opens in new tab) for all extras, and Multimodal Embedding & Reranker Models (opens in new tab) for multimodal support in general):
pip install -U "sentence-transformers[image]"Sentence Transformers v6.0 requires
transformersv5.x,torch2.2+, andhuggingface-hubv1.x. If you pin any of those lower, plan the upgrade first. See the Migration Guide (opens in new tab) for the full list of breaking changes.
Loading a Model
Loading a multi-vector model looks exactly like loading any other Sentence Transformers model:
from sentence_transformers import MultiVectorEncoder
model = MultiVectorEncoder("lightonai/LateOn")
To find models that work, look for the multi-vector and sentence-transformers tags (opens in new tab) on the Hub. Any model with those tags loads with the line above, whether it started life as a PyLate checkpoint, a Stanford-NLP ColBERT checkpoint, or a ColPali-family model for visual document retrieval. We’re working through the ecosystem to get that tag onto every model that works, so the list keeps growing.
Underneath, MultiVectorEncoder reads each of the formats these checkpoints have been published in over the years, so PyLate and Stanford-NLP checkpoints load directly even where the tag hasn’t been added yet:
from sentence_transformers import MultiVectorEncoder
# Native Sentence Transformers checkpoints. PyLate builds on the same schema,
# so any PyLate checkpoint loads identically
model = MultiVectorEncoder("lightonai/LateOn")
model = MultiVectorEncoder("mixedbread-ai/mxbai-edge-colbert-v0-17m")
model = MultiVectorEncoder("LiquidAI/LFM2.5-ColBERT-350M", trust_remote_code=True)
# Any Stanford-NLP ColBERT checkpoint, detected via the \`HF_ColBERT\` architecture
# marker. The inline projection weight and the recipe come from \`artifact.metadata\`
model = MultiVectorEncoder("colbert-ir/colbertv2.0")
model = MultiVectorEncoder("answerdotai/answerai-colbert-small-v1")
# A bare transformer: a fresh random projection is appended, so training is required
model = MultiVectorEncoder("answerdotai/ModernBERT-base")Visual document retrieval models are the exception. ColPali-family checkpoints ship in colpali-engine’s own format, which carries no information Sentence Transformers can use, so each one needs a small configuration added to its repository before it loads. Most of that work is done and waiting to be merged. See Supported Models for the current state and how to load them today.
Inspecting What a Checkpoint Configured
Multi-vector models carry a handful of recipe knobs that differ per checkpoint: marker prefixes for queries and documents, length caps, whether queries are padded out with [MASK] tokens, and which tokens are skipped when scoring documents. All of them live in the module configs, so print(model) shows you exactly what you loaded. Here’s the original ColBERTv2 checkpoint, which pads every query to exactly 32 tokens and truncates documents at 180:
from sentence_transformers import MultiVectorEncoder
model = MultiVectorEncoder("colbert-ir/colbertv2.0")
print(model)
"""
MultiVectorEncoder(
(0): Transformer({..., 'document_length': 180,
'query_expansion': {'strategy': 'fixed', 'attend': False, 'token': None, 'length': 32}})
(1): Dense({'in_features': 768, 'out_features': 128, 'bias': False, ...})
(2): MultiVectorMask({'skiplist_words': ['!', '"', '#', ...], 'skiplist_tasks': ['document'], ...})
(3): Normalize({...})
)
"""
print(model.prompts)
# {'query': '[unused0] ', 'document': '[unused1] '}
That’s the classic ColBERT pipeline: a Transformer producing contextualized token embeddings, a token-level Dense projecting each of them to 128 dimensions, a MultiVectorMask deciding which tokens count during scoring, and a token-level Normalize. Other checkpoints fill in different values. lightonai/GTE-ModernColBERT-v1 uses the same four modules with [Q] and [D] prompts, no query expansion, and caps of 48 and 300.
You rarely need to touch any of this, since every released checkpoint configures its own. It matters when you build a model from a bare backbone, which is covered in Creating Custom Models (opens in new tab).
One value is worth checking against your own data, though. document_length truncates, so anything past it never reaches the index. For example, a 662-token passage through LateOn’s cap of 300 comes back as 273 vectors, with the rest of the passage simply gone. Most of these checkpoints were trained on short passages, so if your chunks are longer than the cap, you can lift it for a single call with encode_document(..., processing_kwargs={"text": {"max_length": 512}}), keeping in mind that you would be running the model past the length it was trained on and that the index grows roughly in proportion. Multi-vector models tend to tolerate that well. On MLDR (opens in new tab), a long-document retrieval benchmark, the multilingual siblings of the pair above show the gap clearly: mLateOn scores 77.92 against mDenseOn’s 51.59 (opens in new tab).
Encoding Queries and Documents
Multi-vector models are asymmetric: queries and documents go through different prefixes, different length caps, and different scoring masks. Unlike many dense models, where the two are interchangeable, encode_query() (opens in new tab) and encode_document() (opens in new tab) are required to get correct embeddings:
from sentence_transformers import MultiVectorEncoder
model = MultiVectorEncoder("lightonai/mLateOn")
queries = ["What is the capital of France?"]
documents = [
"Paris is the capital of France.",
"Berlin is the capital and largest city of Germany, by both area and population.",
]
query_embeddings = model.encode_query(queries)
document_embeddings = model.encode_document(documents)
print(query_embeddings[0].shape)
# (10, 128)
print(document_embeddings[0].shape, document_embeddings[1].shape)
# (10, 128) (19, 128)
Note what you get back: a list of 2D tensors, one per input, each of shape (num_tokens, embedding_dim). Unlike dense embeddings, you can’t stack these into one rectangular tensor, because every input has its own token count. The second document is longer than the first, so it comes back as a taller matrix.
Each call applies the model’s own recipe for you. encode_query prepends the query marker, expands the query to a fixed length if the checkpoint asks for it, and caps it at the query length. encode_document prepends the document marker, caps at the document length, and drops any skiplisted tokens (punctuation, for most checkpoints) from the scoring mask.
The usual encode() arguments all still apply, so batch_size, show_progress_bar, convert_to_tensor, device, and multi-process pools work the way you’d expect:
document_embeddings = model.encode_document(
documents,
batch_size=64,
convert_to_tensor=True,
show_progress_bar=True,
)Scoring with MaxSim
model.similarity() (opens in new tab) computes the full all-pairs MaxSim matrix:
from sentence_transformers import MultiVectorEncoder
model = MultiVectorEncoder("lightonai/LateOn")
query_embeddings = model.encode_query(["Which planet is known as the Red Planet?"])
document_embeddings = model.encode_document([
"Venus is often called Earth's twin because of its similar size and proximity.",
"Mars, known for its reddish appearance, is often referred to as the Red Planet.",
"Jupiter, the largest planet in our solar system, has a prominent red spot.",
"Saturn, famous for its rings, is sometimes mistaken for the Red Planet.",
])
scores = model.similarity(query_embeddings, document_embeddings)
print(scores)
# tensor([[10.7942, 11.1104, 10.9743, 11.0811]])Mars wins, as it should. Note how close the runners-up are: Saturn also contains the literal phrase “the Red Planet”, and Jupiter is a planet with a red spot, so a token-level operator has plenty to latch onto in all three. The ordering is what matters.
Scores often sit this close together, as GLInt (opens in new tab) shows by measuring the spread across a full candidate pool. MaxSim takes a maximum per query token, so a document will usually give every query token some decent best match, and scores start from a floor. Contextualized token embeddings are also anisotropic, clustering in a narrow cone rather than spreading out, so even arbitrary token pairs tend to score high.
There is also model.similarity_pairwise() (opens in new tab), for when you already have matched pairs and just want the pair scores instead of the full similarity matrix:
scores = model.similarity_pairwise(query_embeddings, document_embeddings[:1])
print(scores)
# tensor([10.7942])Score Magnitude and MeanMaxSim
MaxSim sums over query tokens, so its magnitude scales with how many query tokens there are, which means you can’t compare scores across models with different query recipes. LateOn encodes the Red Planet query above as 12 tokens. Run that same query and those same documents through ColBERTv2, which pads and truncates every query to exactly 32 tokens, and the scores land in a completely different range:
model = MultiVectorEncoder("colbert-ir/colbertv2.0")
# ... same encode_query / encode_document / similarity calls ...
print(scores)
# tensor([[12.7970, 27.1945, 23.8495, 24.5656]])Within one model the ordering is all you need, but if you want scores on a bounded scale, switch the model’s similarity function to MeanMaxSim, which divides by the query token count. Back on LateOn:
model = MultiVectorEncoder("lightonai/LateOn", similarity_fn_name="meanmaxsim")
# or on an already-loaded model: model.similarity_fn_name = "meanmaxsim"
print(model.similarity(query_embeddings, document_embeddings))
# tensor([[0.8995, 0.9259, 0.9145, 0.9234]])
Now every score is an average cosine similarity in [-1, 1], although you’ll only see [0, 1] in practice.
Semantic Search
If your corpus is small, exhaustive MaxSim over all of it is the simplest thing that works. Encode the corpus once, then score each query against everything:
import time
from datasets import load_dataset
from sentence_transformers import MultiVectorEncoder
dataset = load_dataset("sentence-transformers/natural-questions", split="train[:5000]")
# Several questions share an answer passage, so drop repeats but keep the order
corpus = list(dict.fromkeys(dataset["answer"])) # 5,000 rows -> 4,874 passages
model = MultiVectorEncoder("lightonai/LateOn")
corpus_embeddings = model.encode_document(corpus, convert_to_tensor=True, show_progress_bar=True)
query = "when did richmond last play in a preliminary final"
start = time.perf_counter()
query_embeddings = model.encode_query([query], convert_to_tensor=True)
scores = model.similarity(query_embeddings, corpus_embeddings)[0] # 98ms
top_scores, top_indices = scores.topk(3)
print(f"Search took .1
ms")
for score, index in zip(top_scores.tolist(), top_indices.tolist()):
print(f"{score:.4f} {corpus[index][:100]}")
"""
Search took 122.7ms
11.9192 Richmond Football Club Richmond began 2017 with 5 straight wins, a feat it had not achieved
11.7591 2017 AFL Grand Final The 2017 AFL Grand Final was an Australian rules football game contest
11.6710 Battle of Appomattox Court House The Battle of Appomattox Court House (Virginia, U.S.), fou
"""Those 4,874 passages encoded in 20 seconds on an RTX 3090, and each search takes about 120ms end to end, most of that the MaxSim scoring against all 608,414 token vectors. This is exact, but it scales linearly in total corpus tokens and keeps every token vector in memory, so reach for it when you have a few thousand documents rather than a few million. The runnable version of this script is semantic_search.py (opens in new tab).
Past that size you want a real late-interaction index, which Sentence Transformers doesn’t ship. It doesn’t need to: these indexes store whatever encode_document produced, so you encode here and hand the token embeddings to something built for them. Indexing has working snippets for four of the options, and the section directly below covers how to skip the index entirely.
Retrieve and Rerank
You can also get late-interaction quality without maintaining a late-interaction index, by using a multi-vector model as your reranker. A fast bi-encoder narrows a large corpus to a handful of candidates, then the multi-vector model rescores only those:
from datasets import load_dataset
from sentence_transformers import MultiVectorEncoder, SentenceTransformer
from sentence_transformers.util import semantic_search
dataset = load_dataset("sentence-transformers/natural-questions", split="train[:50000]")
corpus = list(dict.fromkeys(dataset["answer"]))
retriever = SentenceTransformer("jinaai/jina-embeddings-v5-text-nano-retrieval")
reranker = MultiVectorEncoder("perplexity-ai/pplx-embed-v1-late-0.6b", trust_remote_code=True)
# First stage: index the corpus once with a fast bi-encoder
corpus_embeddings = retriever.encode_document(corpus, convert_to_tensor=True, show_progress_bar=True)
# Retrieve the top 50
query = "when did richmond last play in a preliminary final"
hits = semantic_search(retriever.encode_query([query], convert_to_tensor=True), corpus_embeddings, top_k=50)[0]
candidates = [corpus[hit["corpus_id"]] for hit in hits]
# Second stage: rescore just those candidates with MaxSim
query_embeddings = reranker.encode_query([query])
document_embeddings = reranker.encode_document(candidates)
scores = reranker.similarity(query_embeddings, document_embeddings)[0]
for index in scores.argsort(descending=True)[:3].tolist():
print(f"{scores[index].item():.4f} {candidates[index][:100]}")Only the 50 candidates are ever encoded as multi-vectors, so your index stays a normal dense index and the token vectors are transient. This is the same role a cross-encoder plays in a retrieve-and-rerank stack, but a multi-vector model is considerably cheaper per candidate. You encode the documents in one batch and score them with a matrix multiplication, instead of one forward pass per query-document pair. The runnable script is retrieve_rerank.py (opens in new tab), which prints the timings of both stages.
Indexing
Several vector databases index and score multi-vectors natively: Qdrant (opens in new tab) since v1.10, Weaviate (opens in new tab) since v1.29, Vespa (opens in new tab) for years now, LanceDB (opens in new tab) since v0.15.0, and VectorChord (opens in new tab), which adds a MaxSim operator to Postgres that plain pgvector doesn’t have. Milvus (opens in new tab) joined them in v2.6.4, under array-of-structs rather than the unrelated feature it calls multi-vector search. If you would rather not run a server at all, LightOn’s fast-plaid (opens in new tab) is a pip install away and implements PLAID directly, and PyLate (opens in new tab) wraps it in a fuller retrieval stack.
A few others get you partway. OpenSearch (opens in new tab) and Elasticsearch (opens in new tab) can rescore candidates with MaxSim but not retrieve on it, and the Elasticsearch field is additionally in technical preview and Enterprise-tier. turbopuffer (opens in new tab) has late-interaction indexing in private beta.
The snippets below index text, but nothing in them is text-specific. encode_document hands back the same list of token-vector matrices whether the document was a passage, a page image, an audio clip, or a video, so the ColPali-style models from Visual Document Retrieval go into any of these unchanged. There are simply more vectors per document, which is what makes Token Pooling worth reaching for sooner there.
fast-plaid, Qdrant, Weaviate, and Vespa all take exactly what encode_document returns, so the code is the same up to the client library. Here’s a working snippet for each, run against the 4,874 passages and 608,414 token vectors from the Semantic Search example. Each one carries the ingestion and query times it produced on one machine (RTX 3090, i7-13700K), with no tuning beyond what the code shows, to give a sense of the shape of the work. All four answer the query faster than the 98ms model.similarity took in that section, and three of them do it on the CPU, since fast-plaid is the only one here using the GPU.
All four returned the same three passages in the same order as the exhaustive PyTorch MaxSim earlier in this post, and the three databases reproduce its scores to four decimals! That is because their snippets score every document, which is affordable at this size and removes approximation as a variable. fast-plaid is approximate by design, so its scores differ slightly. The notes under each one say what changes when you switch to an approximate index, which is where rankings start to drift.
fast-plaid
fast-plaid (opens in new tab) is LightOn’s Rust implementation of PLAID, the index ColBERT was originally built around. There’s no server to start, and it reads the tensors encode_document hands back without any conversion.
# pip install sentence-transformers datasets fast-plaid
from datasets import load_dataset
from fast_plaid import search
from sentence_transformers import MultiVectorEncoder
dataset = load_dataset("sentence-transformers/natural-questions", split="train[:5000]")
corpus = list(dict.fromkeys(dataset["answer"]))
model = MultiVectorEncoder("lightonai/LateOn")
query = "when did richmond last play in a preliminary final"
document_embeddings = model.encode_document(corpus, batch_size=32, convert_to_tensor=True)
query_embedding = model.encode_query(query, convert_to_tensor=True)
fast_plaid = search.FastPlaid(index="natural-questions", device="cuda")
# 4,874 documents (608,414 token vectors) indexed in 5s
fast_plaid.create(documents_embeddings=document_embeddings)
results = fast_plaid.search(queries_embeddings=query_embedding.unsqueeze(0), top_k=3) # 11ms
for index, score in results[0]:
print(f"{score:.4f} {corpus[index][:90]}")
"""
11.8828 Richmond Football Club Richmond began 2017 with 5 straight wins, a feat it had not achieve
11.7676 2017 AFL Grand Final The 2017 AFL Grand Final was an Australian rules football game contes
11.6758 Battle of Appomattox Court House The Battle of Appomattox Court House (Virginia, U.S.), fo
"""
The index argument is a directory, not just a label, so the index is written to disk as it is built. Pointing a new FastPlaid at the same path reopens it for searching or for adding more documents, instead of rebuilding from the embeddings each time. On this corpus it occupies 92 MB, against 311.5 MB for the raw float32 vectors.
This is the only one of the four that is approximate, and it is the one place in this section where the scores do not match the exhaustive MaxSim. PLAID prunes with centroids and stores quantized residuals, so the three scores drift by a few hundredths in both directions against the 11.9192 / 11.7591 / 11.6710 computed earlier. The ranking is unaffected here, and that is the trade PLAID is making: it was designed for corpora far larger than this one, where scanning everything is not an option.
Qdrant
Qdrant (opens in new tab) needs a server: docker run -p 6333:6333 qdrant/qdrant. The client also has a local mode (QdrantClient(":memory:")) that needs no server, but it’s a pure-Python reimplementation, so use it for trying things out rather than for timing them.
# pip install sentence-transformers datasets qdrant-client
from datasets import load_dataset
from qdrant_client import QdrantClient, models
from sentence_transformers import MultiVectorEncoder
dataset = load_dataset("sentence-transformers/natural-questions", split="train[:5000]")
corpus = list(dict.fromkeys(dataset["answer"]))
model = MultiVectorEncoder("lightonai/LateOn")
query = "when did richmond last play in a preliminary final"
document_embeddings = model.encode_document(corpus, batch_size=32)
query_embedding = model.encode_query(query)
client = QdrantClient("http://localhost:6333")
client.create_collection(
collection_name="natural-questions",
vectors_config=models.VectorParams(
size=model.get_embedding_dimension(),
distance=models.Distance.COSINE,
multivector_config=models.MultiVectorConfig(
comparator=models.MultiVectorComparator.MAX_SIM
),
# MaxSim never walks the HNSW graph, so skip building one
hnsw_config=models.HnswConfigDiff(m=0),
),
)
# 4,874 documents (608,414 token vectors) ingested in 26.3s
client.upload_points(
collection_name="natural-questions",
points=[
models.PointStruct(id=idx, vector=embedding, payload={"text": text})
for idx, (embedding, text) in enumerate(zip(document_embeddings, corpus))
],
batch_size=64,
)
results = client.query_points(
collection_name="natural-questions",
query=query_embedding,
limit=3,
with_payload=True,
).points # 18ms
for result in results:
print(f"{result.score:.4f} {result.payload['text'][:90]}")
"""
11.9192 Richmond Football Club Richmond began 2017 with 5 straight wins, a feat it had not achieve
11.7591 2017 AFL Grand Final The 2017 AFL Grand Final was an Australian rules football game contes
11.6710 Battle of Appomattox Court House The Battle of Appomattox Court House (Virginia, U.S.), fo
"""MAX_SIM is the only comparator Qdrant offers, and hnsw_config=HnswConfigDiff(m=0) is their recommendation for late-interaction fields, since the vectors are used for rescoring rather than graph traversal. Note that Qdrant themselves suggest reserving late interaction for reranking a few hundred candidates rather than scanning a whole collection, which is the Retrieve and Rerank pattern. At 4,874 documents the full scan costs 18ms and is exact, but that doesn’t extrapolate.
Weaviate
Weaviate (opens in new tab) needs a server too: docker run -p 8080:8080 -p 50051:50051 cr.weaviate.io/semitechnologies/weaviate:1.34.0. Multi-vector support needs 1.29 or newer, and the embedded mode isn’t available on Windows.
# pip install sentence-transformers datasets weaviate-client
import weaviate
from datasets import load_dataset
from sentence_transformers import MultiVectorEncoder
from weaviate.classes.config import Configure, DataType, Property
from weaviate.classes.query import MetadataQuery
dataset = load_dataset("sentence-transformers/natural-questions", split="train[:5000]")
corpus = list(dict.fromkeys(dataset["answer"]))
model = MultiVectorEncoder("lightonai/LateOn")
query = "when did richmond last play in a preliminary final"
document_embeddings = model.encode_document(corpus, batch_size=32)
query_embedding = model.encode_query(query)
client = weaviate.connect_to_local()
collection = client.collections.create(
"Documents",
# self_provided turns on MaxSim late interaction
vector_config=[Configure.MultiVectors.self_provided(name="colbert")],
properties=[Property(name="text", data_type=DataType.TEXT)],
)
# 4,874 documents (608,414 token vectors) ingested in 41s
with collection.batch.fixed_size(batch_size=64) as batch:
for text, embedding in zip(corpus, document_embeddings):
batch.add_object(properties={"text": text}, vector={"colbert": embedding.tolist()})
results = collection.query.near_vector(
near_vector=query_embedding.tolist(),
target_vector="colbert",
limit=3,
return_metadata=MetadataQuery(distance=True),
) # 17ms
for result in results.objects:
# Weaviate reports the MaxSim score as a negated distance
print(f"{-result.metadata.distance:.4f} {result.properties['text'][:90]}")
"""
11.9192 Richmond Football Club Richmond began 2017 with 5 straight wins, a feat it had not achieve
11.7591 2017 AFL Grand Final The 2017 AFL Grand Final was an Australian rules football game contes
11.6710 Battle of Appomattox Court House The Battle of Appomattox Court House (Virginia, U.S.), fo
"""
client.close()
Defaults are enough here: Weaviate’s dynamic ef resolves to 100 for a top-3 query, and this ranking is already exact from about 32 upward. That margin is a property of the embeddings rather than of Weaviate, so it’s worth confirming on your own model instead of assuming the defaults hold.
Weaviate also supports MUVERA encoding, which made ingestion 3x faster and queries 1.8x faster in our test. It cost far more accuracy than that speed is worth at this size though: the correct third passage didn’t appear even in its top 50.
Vespa
Vespa (opens in new tab) also runs in a container, but pyvespa starts it for you, so there’s no separate docker run.
# pip install sentence-transformers datasets pyvespa
from datasets import load_dataset
from sentence_transformers import MultiVectorEncoder
from vespa.deployment import VespaDocker
from vespa.package import (
ApplicationPackage, Document, Field, FirstPhaseRanking, Function, RankProfile, Schema,
)
dataset = load_dataset("sentence-transformers/natural-questions", split="train[:5000]")
corpus = list(dict.fromkeys(dataset["answer"]))
model = MultiVectorEncoder("lightonai/LateOn")
query = "when did richmond last play in a preliminary final"
document_embeddings = model.encode_document(corpus, batch_size=32)
query_embedding = model.encode_query(query)
# "dt" is a mapped dimension over the variable token count, "x" the dense 128-dim vector
package = ApplicationPackage(
name="colbert",
schema=[
Schema(
name="doc",
document=Document(fields=[
Field(name="text", type="string", indexing=["summary"]),
Field(name="colbert", type="tensor<float>(dt{}, x[128])", indexing=["attribute"]),
]),
rank_profiles=[
RankProfile(
name="colbert",
inputs=[("query(qt)", "tensor<float>(qt{}, x[128])")],
functions=[Function(
name="max_sim", # per query token take the best document token, then sum
expression="sum(reduce(sum(query(qt) * attribute(colbert), x), max, dt), qt)",
)],
first_phase=FirstPhaseRanking(expression="max_sim"),
)
],
)
],
)
app = VespaDocker(port=8080).deploy(application_package=package) # ~40s to boot
# Vespa reads a mixed tensor as {token index: vector}, for documents and queries alike
def to_tensor(embedding):
return {str(token): vector for token, vector in enumerate(embedding.tolist())}
# 4,874 documents (608,414 token vectors) ingested in ~80s
app.feed_iterable(
({"id": str(idx), "fields": {"text": text, "colbert": to_tensor(embedding)}}
for idx, (text, embedding) in enumerate(zip(corpus, document_embeddings))),
schema="doc",
)
response = app.query(body={
"yql": "select text from doc where true",
"ranking.profile": "colbert",
"hits": 3,
"input.query(qt)": to_tensor(query_embedding),
}) # ~75ms warm, ~115ms on the first call
for hit in response.hits:
print(f"{hit['relevance']:.4f} {hit['fields']['text'][:90]}")
"""
11.9192 Richmond Football Club Richmond began 2017 with 5 straight wins, a feat it had not achieve
11.7591 2017 AFL Grand Final The 2017 AFL Grand Final was an Australian rules football game contes
11.6710 Battle of Appomattox Court House The Battle of Appomattox Court House (Virginia, U.S.), fo
"""
Vespa asks for the most upfront structure of the four, because you’re declaring a ranking pipeline rather than just an index. In exchange you get to write MaxSim out as a tensor expression and see exactly what it computes. This version puts MaxSim in first-phase over where true, which scores all 4,874 documents and is why the output matches exhaustive MaxSim exactly. It’s deliberately not what Vespa recommends at scale: their ColBERT sample app (opens in new tab) stores int8-binarized vectors and moves MaxSim into second-phase to rerank a cheaper first stage.
Moving to that phased setup needs care: second-phase rescores only the best 100 candidates by default, and here that window left two of the three correct passages unscored entirely. Raising rerank-count to cover your candidate set fixes that, though at this size the phased version still came out slower than simply scanning everything.
Visual Document Retrieval
Late interaction is the state of the art for visual document retrieval: matching a text query against page images, with charts, tables, and layout intact, and no OCR step. This is what the ColPali (opens in new tab) family of models does, and those checkpoints load and run through the same API, with the revision pinning the open pull request that adds this one’s Sentence Transformers configuration (Supported Models has the full list). Image documents are passed as URLs, local paths, or PIL images:
from sentence_transformers import MultiVectorEncoder
model = MultiVectorEncoder("vidore/colqwen2.5-v0.2")
queries = [
"What is the variable represented on the y-axis of the graph?",
"Total outlay is maximum in which year?",
]
images = [
"https://huggingface.co/datasets/sentence-transformers/example-documents/resolve/main/doc1.jpg",
"https://huggingface.co/datasets/sentence-transformers/example-documents/resolve/main/doc2.jpg",
"https://huggingface.co/datasets/sentence-transformers/example-documents/resolve/main/doc3.jpg",
"https://huggingface.co/datasets/sentence-transformers/example-documents/resolve/main/doc4.jpg",
]
query_embeddings = model.encode_query(queries)
document_embeddings = model.encode_document(images)
print(query_embeddings[0].shape, document_embeddings[0].shape)
# (25, 128) (755, 128)
scores = model.similarity(query_embeddings, document_embeddings)
print(scores)
# tensor([[13.8672, 12.3115, 12.1670, 11.0293],
# [ 7.2012, 14.7207, 6.9414, 6.9746]])Each query retrieves its own page (the diagonal), and the second query separates much more cleanly than the first, since only one of the four pages is about outlay over time.
The code is unchanged. Underneath, the processor handles the visual prompt and the image patches, and MaxSim scores query text tokens against document image patches. A page holds many separate regions, which is exactly what makes late interaction a natural fit here, since a single vector would have to average a chart, a table, and three paragraphs into one summary. That fidelity costs index space, though. The shapes above are 755 token vectors for one page against 25 for the query, where a Natural Questions passage from earlier averaged about 125, so token pooling is worth reaching for earlier here than it is for text.
These are VLMs, so plan for the memory they need. The table in Supported Models runs from 252M to 8.8B parameters, and the small end of it stays practical on CPU where the multi-billion ones don’t.
Page images are the common case, but they’re not the only non-text modality. Sentence Transformers accepts text, images, audio, and video, and a checkpoint supports whichever of those its processor does, which model.modalities reports. A single document can combine modalities too, by passing a dict like {"text": ..., "image": ...} in place of a bare value. Multimodal Embedding & Reranker Models (opens in new tab) covers multimodal models in Sentence Transformers more broadly, and the Usage documentation (opens in new tab) lists exactly which input formats each modality accepts.
Audio Retrieval
vidore/colqwen-omni-v0.1 (opens in new tab) is built on Qwen2.5-Omni and takes all four modalities. Retrieving a recorded conversation with it is the same two calls as retrieving a page:
# pip install -U "sentence-transformers[audio,video]"
import torch
from datasets import Audio, load_dataset
from sentence_transformers import MultiVectorEncoder
model = MultiVectorEncoder(
"vidore/colqwen-omni-v0.1",
model_kwargs={"dtype": torch.bfloat16},
)
print(model.modalities)
# ['text', 'image', 'audio', 'video', 'message']
# 20 recorded conversations, averaging 28 seconds each
dataset = load_dataset("eustlb/dailytalk-conversations-grouped", split="train[:20]")
dataset = dataset.cast_column("audio", Audio(sampling_rate=16_000))
audio = [row["array"] for row in dataset["audio"]] # raw mono waveforms, float32 at 16 kHz
query_embeddings = model.encode_query(["medicine for car nausea"])
document_embeddings = model.encode_document(audio, batch_size=2)
scores = model.similarity(query_embeddings, document_embeddings)[0]
top_scores, top_indices = scores.topk(3)
for score, index in zip(top_scores.tolist(), top_indices.tolist()):
print(f"{score:.4f} {' / '.join(dataset[index]['texts'][:2])}")
"""
50.8902 Excuse me? Do you have anything for a carsickness? / Yes, but you look fine.
46.1028 Excuse me, could you tell me where you have got that music book? / Certainly. Let me see. Oh, it's on that shelf.
46.0514 Jeff, I'm going to the supermarket. Do you want to come with me? / I think the supermarket is closed now.
"""ColQwen-Omni was trained purely on image-text pairs, so its audio retrieval is zero-shot: it never heard a training example, and there is no transcription step anywhere in the pipeline. The query says nausea where the recording says carsickness, and it still picks the pharmacy conversation out of twenty by a wide margin.
Video Retrieval
Video works the same way, but sample the frames or it will eat your VRAM. Its release blogpost (opens in new tab) is blunt about this, that video “is very memory-intensive, so it’s best suited for short clips”:
import torch
from sentence_transformers import MultiVectorEncoder
model = MultiVectorEncoder(
"vidore/colqwen-omni-v0.1",
model_kwargs={"dtype": torch.bfloat16},
)
# Sparse, low-resolution frames: 0.5 fps rather than the full frame rate
model[0].processing_kwargs.update(
{"video": {"max_pixels": 32 * 28 * 28, "do_sample_frames": True, "fps": 0.5}}
)
query_embeddings = model.encode_query(["How to cook Mapo Tofu?"])
document_embeddings = model.encode_document([
"https://huggingface.co/datasets/sentence-transformers/example-documents/resolve/main/mapo_tofu.mp4",
"https://huggingface.co/datasets/sentence-transformers/example-documents/resolve/main/zhajiang_noodle.mp4",
], batch_size=1)
print(model.similarity(query_embeddings, document_embeddings))
# tensor([[53.3100, 51.0561]])At 1 fps and full resolution the same pair of videos produces 8,426 and 5,137 token vectors and peaks at 20.8 GB of VRAM, against 4,240 and 2,446 vectors and 12.5 GB here, for a model that occupies 9.0 GB on its own. The ranking is identical either way. Long audio wants the same treatment, and the release blogpost recommends 30-second chunks, which come to roughly 800 tokens each.
Interpretability
Because MaxSim is a sum of per-query-token maxima, a ranking decomposes exactly: every point of a document’s score belongs to one query token and one document token. That lets you answer “why did this rank here?” precisely, rather than by eye.
For image documents, sentence_transformers.multi_vector_encoder.interpretability overlays that decomposition onto the page as the standard ColPali heatmap, either aggregated over the query or one map per query token. Asking “How much was spent on water resources and power?” against the outlays page from above, this is where the water token went:
heatmap.py (opens in new tab) is the runnable version, including the masking step that lines the document embedding up with the patch grid.
Text documents have no patch grid to overlay, but the same decomposition applies. text_similarity_map.py (opens in new tab) ranks a corpus and then attributes the top hit’s score token by token, here on the Natural Questions corpus from earlier with the 32M-parameter mxbai-edge-colbert-v0-32m (opens in new tab):
Query: when did richmond last play in a preliminary final
Top 3 of 4874 documents by exhaustive MaxSim (191.0ms):
12.3489 Richmond Football Club Richmond began 2017 with 5 straight wins, a feat it had not achieved since 19
12.1771 2017 AFL Grand Final The 2017 AFL Grand Final was an Australian rules football game contested betwee
12.0591 2018 UEFA Champions League Final The 2018 UEFA Champions League Final was the final match of the 201
query token best document token sim share
when since 0.9154 7.4%
did had 0.9675 7.8%
rich rich 0.9764 7.9%
mond mond 0.9856 8.0%
last to 0.9249 7.5%
play game 0.9384 7.6%
in the 0.9732 7.9%
a a 0.9587 7.8%
preliminary preliminary 0.9394 7.6%
final final 0.9654 7.8%
--------------------------------------------------------
3 special tokens 2.8038 22.7%
MaxSim score 12.3489 100.0%rich, mond, preliminary, and final matched themselves, while when settled on since and play on game. The special tokens are worth noticing too: three of them contribute 22.7% of the score while carrying none of the query’s content. Below this table the script prints the passage itself, with the winning tokens highlighted in place.
Token Pooling
If the index footprint worries you, the most effective knob is to store fewer token vectors. HierarchicalTokenPooling implements the token pooling (opens in new tab) technique from Clavié, Chaffin, and Adams: it clusters each document’s token vectors with Ward linkage on cosine distance and replaces each cluster with its mean, keeping roughly 1 / pool_factor of the tokens. Within one document a lot of token vectors end up close to each other, so much of what you drop is redundancy rather than signal:
from datasets import load_dataset
from sentence_transformers import MultiVectorEncoder
from sentence_transformers.multi_vector_encoder.modules import HierarchicalTokenPooling
dataset = load_dataset("sentence-transformers/natural-questions", split="train[:5000]")
documents = list(dict.fromkeys(dataset["answer"]))
model = MultiVectorEncoder("lightonai/LateOn")
pooling = HierarchicalTokenPooling(pool_factor=2)
document_embeddings = model.encode_document(documents, token_pooling=pooling)There are three places to apply it, depending on when you want to pay for it:
# 1. Per encode call, as above
document_embeddings = model.encode_document(documents, token_pooling=pooling)
# 2. Standalone, on embeddings you already have saved (e.g. list of [num_tokens, num_dims] tensors)
pooled = pooling.pool(document_embeddings)
# 3. Baked into the model, so every consumer of the checkpoint gets pooled documents
model.append(HierarchicalTokenPooling(pool_factor=2))
model.save_pretrained("my-pooled-colbert")
By default, pooling applies to documents only, since queries are short and are the side you can’t afford to distort. On the Natural Questions corpus from earlier, the reduction tracks pool_factor closely, and pooling all 608k token vectors took about 6 seconds:
pool_factor | Token vectors | Reduction | float32 index |
|---|---|---|---|
| 1 (off) | 608,414 | 1.00x | 311.5 MB |
| 2 | 305,438 | 1.99x | 156.4 MB |
| 3 | 204,407 | 2.98x | 104.7 MB |
| 4 | 153,936 | 3.95x | 78.8 MB |
A cluster mean is a worse match for a query token than the best of its members was, and the coarser the clusters, the more that shows. The original experiments (opens in new tab) measured that cost on BEIR and found very little of it: 100.6% of the unpooled retrieval performance on average at pool_factor=2, and 99.0% at pool_factor=3. Halving your index for free is a good deal, so 2 is a reasonable place to start. How much it costs on your data is corpus-specific though, so measure it with an evaluator before you settle on a factor. The runnable comparison is token_pooling.py (opens in new tab).
How far you can push pool_factor is also partly a property of the model. LightOn’s hierarchical pooling regularization (opens in new tab) trains for exactly that, shaping the embedding space so pooling costs less and reporting 99.4% retention at 5x compression. Training with that regularizer isn’t in Sentence Transformers yet, but the resulting checkpoints are ordinary PyLate models, so lightonai/LateOn-hpool-regularized (opens in new tab) loads and pools like any other.
Speeding Up Inference
Multi-vector models run through the same backend machinery as the rest of Sentence Transformers, so you get torch (default), onnx, and openvino, alongside half precision, Flash Attention, and torch.compile.
On GPU, fp16 with Flash Attention is the best configuration we measured, at 2.44x the throughput of fp32 with no measurable retrieval quality loss. Flash Attention helps multi-vector models more than most, because documents are only truncated and never padded to a shared length, so your batches have widely varying sequence lengths that unpadding can exploit:
from sentence_transformers import MultiVectorEncoder
model = MultiVectorEncoder(
"lightonai/GTE-ModernColBERT-v1",
model_kwargs={"attn_implementation": "flash_attention_2", "dtype": "float16"},
)
GPU

CPU
Models with non-attend query expansion (
attend=False, which covers the Stanford-NLP checkpoints likecolbert-ir/colbertv2.0andanswerdotai/answerai-colbert-small-v1) reject Flash Attention at load time. Flash Attention stripsattention_mask=0positions, so the[MASK]expansion tokens that MaxSim scores would never receive an attention update. Use"sdpa"for those models.
On CPU, OpenVINO is your better bet where the architecture is supported, and int8 quantization buys a further speedup at a cost of about 0.4% accuracy. See Speeding up Inference (opens in new tab) for the full benchmark details, the export and quantization helpers, and a flowchart for picking a backend.
Evaluating a Model
MultiVectorNanoBEIREvaluator runs the NanoBEIR (opens in new tab) suite of 13 small BEIR subsets with MaxSim scoring, and needs no data preparation on your side:
from sentence_transformers import MultiVectorEncoder
from sentence_transformers.multi_vector_encoder.evaluation import MultiVectorNanoBEIREvaluator
model = MultiVectorEncoder("lightonai/GTE-ModernColBERT-v1")
evaluator = MultiVectorNanoBEIREvaluator(batch_size=16)
results = evaluator(model)
print(f"{evaluator.primary_metric}: {results[evaluator.primary_metric]:.4f}")This also makes it easy to check the claim from the top of this post. lightonai/LateOn (opens in new tab) and lightonai/DenseOn (opens in new tab) were trained by LightOn on the same data with the same ModernBERT backbone and the same 149M parameters, differing only in whether they keep one vector per token or pool down to one per document. Running both over all 13 NanoBEIR datasets isolates what that choice buys:
| NanoBEIR dataset | LateOn (multi-vector, 128d) | DenseOn (dense, 768d) |
|---|---|---|
| MSMARCO | 0.7194 | 0.6517 |
| NQ | 0.7810 | 0.7511 |
| HotpotQA | 0.9295 | 0.8802 |
| FEVER | 0.9702 | 0.9612 |
| ClimateFEVER | 0.4887 | 0.4846 |
| DBPedia | 0.6836 | 0.6748 |
| QuoraRetrieval | 0.9795 | 0.9687 |
| Touche2020 | 0.5938 | 0.5673 |
| ArguAna | 0.5562 | 0.5660 |
| NFCorpus | 0.3949 | 0.3851 |
| SciFact | 0.7978 | 0.8057 |
| SCIDOCS | 0.4469 | 0.4484 |
| FiQA2018 | 0.5871 | 0.6491 |
| Mean | 0.6868 | 0.6764 |
Late interaction wins on 9 of the 13 datasets and on the mean, by roughly one NDCG point. The four it loses (ArguAna, FiQA2018, SCIDOCS, and SciFact) are the shape of the tradeoff you should expect: a real gain in retrieval quality at the same model size, paid for in index footprint, rather than a universal win on every dataset. The same pair scores 57.22 against 56.20 on the full 15-dataset BEIR, a comparable gap, so the margin is not an artifact of the small benchmark.
Alongside NanoBEIR, MultiVectorInformationRetrievalEvaluator, MultiVectorRerankingEvaluator, MultiVectorTripletEvaluator, and MultiVectorDistillationEvaluator cover the usual evaluation setups on your own data. They’re documented in the Evaluation API Reference (opens in new tab).
Coming from PyLate or colpali-engine
MultiVectorEncoder absorbs the modeling, inference, training, and evaluation of both libraries. Every PyLate checkpoint loads directly, and Supported Models lists the colpali-engine checkpoints along with the revision to pass where one is still needed. If you’re migrating, these are the calls that change:
| PyLate | Sentence Transformers |
|---|---|
pylate.models.ColBERT(model_name_or_path=...) | MultiVectorEncoder(...) |
model.encode(..., is_query=True) | model.encode_query(...) |
model.encode(..., is_query=False) | model.encode_document(...) |
pylate.scores.colbert_scores | model.similarity |
pylate.indexes.PLAID / pylate.retrieve.ColBERT | no equivalent, keep PyLate’s PLAID or see Indexing |
| colpali-engine | Sentence Transformers |
|---|---|
ColQwen2.from_pretrained(...) + ColQwen2Processor | MultiVectorEncoder(...) |
processor.process_queries(...) + model(**batch) | model.encode_query(queries) |
processor.process_images(...) + model(**batch) | model.encode_document(images) |
processor.score_multi_vector(qs, ds) | model.similarity(query_embeddings, document_embeddings) |
mask_non_image_embeddings=True | MultiVectorMask(keep_only_token_ids=[...]) |
HierarchicalTokenPooler | HierarchicalTokenPooling |
colpali_engine.interpretability | sentence_transformers.multi_vector_encoder.interpretability |
One difference worth calling out: on a bare (non-ColBERT) checkpoint, PyLate’s ColBERT("bert-base-uncased") applies the classic recipe by default, while MultiVectorEncoder("bert-base-uncased") builds a plain stack and leaves the prefixes, query expansion, and skiplist as explicit choices. The training loss and evaluator equivalents, and the data-handling differences, are in the Migration Guide (opens in new tab).
Note that save compatibility is one-way in every case: PyLate, Stanford-NLP ColBERT, and colpali-engine checkpoints all load into MultiVectorEncoder, but MultiVectorEncoder.save_pretrained output isn’t loadable by any of them.
Supported Models
Models carrying the multi-vector and sentence-transformers tags (opens in new tab) on the Hub are the list that stays current, and we’re working to get those tags onto every model that works. The tables below are what we test against directly, so treat them as a starting point rather than the full set. For text retrieval in particular, any PyLate or Stanford-NLP ColBERT checkpoint loads whether or not it carries the tag yet.
Some entries need a small Sentence Transformers configuration added to their repository first, and several of those are still open pull requests at the time of writing. Where a revision is listed below, pass it until that pull request is merged, after which the plain model name is enough:
model = MultiVectorEncoder("vidore/colqwen-omni-v0.1", revision="refs/pr/N")Text Retrieval Models
These load with their trained prefix tokens, query expansion, and punctuation skiplist recovered from the saved configuration.
The NanoBEIR column reports the mean NDCG@10 (higher is better) across the 13 NanoBEIR datasets (opens in new tab), each a 50-query subsample of a BEIR dataset, as a fast proxy for English text retrieval quality. We used the MultiVectorNanoBEIREvaluator to compute the scores for the primarily-English models. A - means the model was not evaluated on it. Note that NanoBEIR is a small benchmark, and its scores aren’t a substitute for evaluating on your own data, which is always the right way to pick a model.
Visual Document Retrieval Models
ColPali-style models embed page images as documents and text as queries.
The NanoViDoRe column reports the mean NDCG@10 (higher is better) across NanoViDoRe v3 (opens in new tab), a compact visual document retrieval benchmark spanning 8 subsets (computer science, energy, finance in English and French, HR, industrial, pharmaceuticals, and physics). Like with NanoBEIR, NanoViDoRe is a small benchmark which shouldn’t replace evaluation on your own data.
Most of these are LoRA adapter repositories, with the adapter applied directly onto its base at load time. Some also have a -merged sibling on the Hub (e.g. vidore/colpali-v1.3-merged (opens in new tab)) with the adapter already folded into the weights.
The three -hf entries are the transformers-native *ForRetrieval ports. They load without any configuration, but use more modeling from transformers and less from sentence_transformers. Generally, it’s preferable to use the original models instead, as the ports score approximately the same.
Acknowledgements
Late interaction in Sentence Transformers rests on a lot of earlier work. Thanks to Omar Khattab and Matei Zaharia for ColBERT (opens in new tab), which everything here descends from, and to the LightOn team (Antoine Chaffin, Raphael Sourty, Paulo Moura, and Amélie Chatelain) for PyLate (opens in new tab) and fast-plaid (opens in new tab), which carried late interaction for years and shaped a good deal of the API described above.
Thanks to the ColPali team (Manuel Faysse, Hugues Sibille, Tony Wu, Bilel Omrani, Gautier Viaud, Céline Hudelot, and Pierre Colombo) for ColPali (opens in new tab) and colpali-engine, which brought late interaction to page images, and to Benjamin Clavié, Antoine Chaffin, and Griffin Adams for token pooling (opens in new tab).
Thanks as well to the core MTEB team, Kenneth Enevoldsen and Roman Solomatin among many others, for MTEB (opens in new tab) and for the kind of hidden work that keeps information retrieval research running.
And thanks to everyone who trained and released the checkpoints in Supported Models. Without them this post would have had nothing to measure.
Additional Resources
Documentation
- Multi-Vector Encoder > Usage (opens in new tab)
- Multi-Vector Encoder > Pretrained Models (opens in new tab)
- Multi-Vector Encoder > Creating Custom Models (opens in new tab)
- Multi-Vector Encoder > Speeding up Inference (opens in new tab)
- Multi-Vector Encoder > API Reference (opens in new tab)
- Installation (opens in new tab)
- Migration Guide (opens in new tab)
Example Scripts
Training
To learn how to train or finetune these models on your own data:
- Multi-Vector Encoder > Training Overview (opens in new tab)
- Multi-Vector Encoder > Loss Overview (opens in new tab)
- Multi-Vector Encoder > Training Examples (opens in new tab)
- LateOn and mLateOn training scripts (opens in new tab): LightOn’s PyLate recipes for LateOn, mLateOn, DenseOn, and mDenseOn, where the finetuning scripts show practical details like splitting a 16,384-example batch into mini-batches of 16.
Hugging Face Hub
Companion Blogposts
- Training and Finetuning Embedding Models with Sentence Transformers (opens in new tab): the general training guide for text-only dense embedding models.
- Training and Finetuning Reranker Models with Sentence Transformers (opens in new tab): Cross Encoder training, the other way to add a precise second stage.
- Training and Finetuning Sparse Embedding Models with Sentence Transformers (opens in new tab): SPLADE and other sparse encoders, which combine well with late interaction in hybrid search.
- Multimodal Embedding & Reranker Models with Sentence Transformers (opens in new tab): single-vector multimodal models, the dense counterpart to ColPali-style retrieval.
- Training and Finetuning Multimodal Embedding & Reranker Models with Sentence Transformers (opens in new tab): includes a Visual Document Retrieval walkthrough with single-vector models.
- 🪆 Introduction to Matryoshka Embedding Models (opens in new tab): shrink dense embeddings by dimension, the way token pooling shrinks multi-vector ones by count.
More Articles from our Blog
multimodalnlpcommunity
79
April 16, 2026
multimodalnlpcommunity
70
April 9, 2026
Sentence Transformers v6.0 (Hugging Face) adds a fourth model type,
MultiVectorEncoder, bringing ColBERT-style late interaction retrieval natively into the library: PyLate checkpoints, Stanford-NLP ColBERT checkpoints, and colpali-engine visual document retrieval models all load through the standard API, with the new class absorbing the modeling, inference, training, and evaluation of both PyLate and colpali-engine . v6.0 requires transformers v5.x, torch 2.2+, and huggingface-hub v1.x .Multi-vector models keep one vector per token (classically 128-dim) and score with MaxSim — each query token's highest cosine against any document token, summed — preserving token-level matches that a dense single vector averages away; the post calls late interaction the state of the art for visual document retrieval (text queries against page images, no OCR) . The cost is index size: LateOn encoding of 4,874 Natural Questions passages produced 608,414 token vectors (311.5 MB float32 vs 7.5 MB dense MiniLM, ~42x), but PLAID compression cuts that to 92 MB — comparable to an 80 MB dense 4096-dim Qwen3-Embedding-8B index .
In a controlled same-backbone comparison (LightOn's LateOn vs DenseOn: ModernBERT, 149M params, same data), late interaction wins 9 of 13 NanoBEIR datasets and the mean — 0.6868 vs 0.6764, roughly one NDCG point — plus 57.22 vs 56.20 on full BEIR . Hierarchical token pooling (clustering token vectors with Ward linkage on cosine distance) cuts vector count ~2x while the original BEIR experiments measured 100.6% of unpooled retrieval performance at pool_factor=2 and 99.0% at 3x; LightOn's hpool-regularized checkpoints report 99.4% retention at 5x compression .
Serving: fp16 + Flash Attention gives 2.44x fp32 throughput with no measurable retrieval quality loss; OpenVINO int8 on CPU buys speed at ~0.4% accuracy; Stanford-NLP-style checkpoints with non-attend query expansion (e.g. ColBERTv2) reject Flash Attention . Native multi-vector indexing is supported by Qdrant (v1.10+), Weaviate (v1.29+), Vespa, LanceDB (v0.15+), VectorChord, and Milvus (v2.6.4); OpenSearch and Elasticsearch support MaxSim rescoring only (Elasticsearch's field is Enterprise-tier technical preview) .
Notable models in the supported rosters: Perplexity's pplx-embed-v1-late-0.6b (596M; NanoBEIR 0.6662), LiquidAI's LFM2.5-ColBERT-350M (0.6864), AnswerAI's answerai-colbert-small-v1 (33M; 0.6550); on visual document retrieval's NanoViDoRe, webAI-ColVec1.1-8b (8.4B) leads at 0.6580 ahead of Tencent's EVIE-Preview-4.5B (0.6405) and TomoroAI's tomoro-colqwen3-embed-8b (8.8B; 0.6206). The multimodal colqwen-omni-v0.1 adds zero-shot audio retrieval — no transcription step, trained only on image-text pairs .

