ZeroNoise Logo zeronoise
Post
Disconfirm the hire; specialize the model
10 hours ago
2 min read
168 docs
A concise resource brief built around two high-signal recommendation clusters: an adversarial reference-checking playbook surfaced by Sarah Tavel and Sarah Guo, and Harvey’s domain-specific post-training research update highlighted by Sarah Guo and Aaron Levie.

The strongest new recommendations form a practical pairing: make hiring decisions more adversarial, then tailor AI systems to the workflows where specialization can pay off.

Standout: How to run a real reference check

  • Content type / creator: X thread by @dittycheria.
  • Link:Read the thread
  • Recommended by: Sarah Tavel called it “so much wisdom” on references, while Sarah Guo called it “great advice for founders on references.”
  • Key takeaway: A reference check should try to disconfirm the hire—not validate a decision that has already hardened. The thread recommends asking what you are missing, where the candidate is weaker than they appear, and what you will wish you had known six months later.
  • Framework: Seek “blind” references from former bosses, peers, direct reports, and customers rather than relying only on candidate-provided names; discount praise and amplify criticism to counter politeness; ask questions that force specific examples; and use forced rankings such as “top 1%,” “top 10%,” or “top 25%.”
  • Why it matters: This turns references into an adversarial decision tool for senior hiring. Its final test is practical: discover why you may be making the wrong choice while you can still change it.

Technical companion: Update on our post-training effort

  • Content type / creator: Technical X article and research update by Harvey researchers @nikogrupen, @ItsJulioPereyra, @calvincongelado, @vtrengarajan, and @gabepereyra.
  • Link:Read the article
  • Recommended by: Sarah Guo described it as a blueprint for the performance and efficiency gains possible from post-training, in-domain data, and workflow understanding. Aaron Levie called it a strong account of how applied-AI companies can lower costs and improve accuracy, while stressing that specialization makes sense only when domain expertise and repeated task volume justify it.
  • Key takeaway: Harvey reports on Tenet, a Kimi K3 base model post-trained with Fireworks for long-horizon legal work, and characterizes its initial results as promising for performance and cost efficiency.
  • Why it matters: The useful decision rule is narrower than “post-train everything”: specialize when a team understands an enterprise workflow deeply and has enough similar work for optimization to matter; otherwise, a general-purpose frontier model may be sufficient.
  • Caveat: This is a first-party Harvey update, so its performance claims are company-reported. Read it primarily for the workflow, evaluation, and training design rather than as independent benchmark validation.
Disconfirm the hire; specialize the model
Patrick OShaughnessy

Patrick O'Shaughnessy recommended Ben Thompson's Stratechery: Thompson has written it for over a decade and O'Shaughnessy calls him "one of my favorite business thinkers" . He shared his video conversation with Thompson (linked at https://x.com/patrick_oshag/status/2089713931183153293), covering why it could be problematic for the US to win the AI race, whether AI funding will run out, Google becoming Berkshire Hathaway, why ads are amazing, TSMC/Intel/Samsung, Nvidia's invisible price cuts and biggest competitors, and Microsoft, Amazon, Apple, and Meta .

In a clip from that conversation, Thompson argues Silicon Valley relearns every ~10 years that consumers do not want to pay for software and do not care about being productive, so consumer businesses must make money from advertising; Dropbox is the canonical example (it had to rebuild and sell to companies), and OpenAI is "replaying the Dropbox story at 100x the size" — it sold many consumer subscriptions but not enough, and would have a great ad product today if it had leaned into advertising as soon as ChatGPT was a hit .

My conversation with [@benthompson](https://x.com/benthompson). Ben has been writing Stratechery for over a decade and remains one of my … Ben Thompson on the two things Silicon Valley relearns every 10 years: 1) "Consumers do not want to pay for software" 2) "Consumers do no…
sarah guo

Sarah Guo (@saranormous) recommended Harvey's X article "Update on our post-training effort" by @gabepereyra and the Harvey research team, calling it "a clear blueprint for the huge performance and efficiency gains possible from post training and in-domain data and workflow understanding" and adding "own your intelligence!" . Article: https://x.com/i/article/2090114065729503232. The article covers Harvey's post-training of a Kimi K3 base into a model called Tenet (with Fireworks), RL training via GSPO, and initial results across M&A diligence, review tables, and firm knowledge .

.@gabepereyra, the research team at [@harvey](https://x.com/harvey), and their partners are giving everyone building specialized intellig… [https://x.com/i/article/2090114065729503232](https://x.com/i/article/2090114065729503232) Update on our post-training effort
sarah guo

Sarah Guo (@saranormous) recommended @dittycheria's Twitter thread on reference checking, calling it 'Great advice for founders on references' and linking to it . The thread — from a founder whose successful company does full reference checks — argues reference checks should aim to disconfirm a candidate, not validate a decision . Key tactics: seek blind references (former bosses, peers, direct reports, customers) rather than relying on candidate-supplied names ; discount praise ~3x and amplify criticism ~3x to offset politeness bias ; attend to qualifications and hesitations more than compliments ; ask questions that force specifics (direct reports' hardest part, repeated feedback, biggest blind spot, behavior in crises, likely failure reason, would you hire again?) ; and use forced ranking (top 1%/10%/25%) to force a real judgment . The core message: use reference checks to find why you might be wrong while you can still change the decision . Link: https://x.com/dittycheria/status/2090414261545640095

Great advice for founders on references ⬇️ [https://x.com/dittycheria/status/2090414261545640095](https://x.com/dittycheria/status/2090414… Yesterday, the founder of a well-known, successful company called me to reference check two people he was considering for a senior role. …
Patrick OShaughnessy

Patrick OShaughnessy recommends Ben Thompson's Stratechery: Thompson has written it for over a decade and is "one of my favorite business thinkers"; OShaughnessy values their conversations on markets and technology .

My conversation with [@benthompson](https://x.com/benthompson). Ben has been writing Stratechery for over a decade and remains one of my …
20VC with Harry Stebbings

During a discussion of SpaceX's $60B all-stock acquisition of Cursor , a 20VC host recommended a blog post by Noah Smith, written about a year before the episode aired (Aug 2026), calling it "a great piece" that argues "only a fool denies that Elon Musk is wildly effective." The host noted Smith is "a moderate centrist blogger who's not an Elon fan," which strengthened the endorsement, and cited the piece to support the view that Musk is one of the most effective people on the planet at getting things done in industrialization, physical AI, and AI . Type: article/blog post; author: Noah Smith; no article title or URL given.

Stripe's $8B OpenRouter Bet | Anthropic's First Profit & The Math Behind Reaching $600B in Revenue?
Aaron Levie

Aaron Levie recommended Harvey's X article 'Update on our post-training effort' as a great post on what post-training looks like for applied AI use-cases, saying it is a compelling value proposition for applied AI companies . The article, by Harvey researchers @nikogrupen, @ItsJulioPereyra, @calvincongelado, @vtrengarajan, and @gabepereyra, describes Tenet, a Kimi K3 base post-trained with Fireworks for long-horizon legal work . Levie's key takeaway: once a company understands a domain deeply and has enough volume of similar tasks, it makes sense to purpose-design models for that work, bringing down costs and improving accuracy . He highlighted the post's point that reward shaping incentivized efficient tool use and reasoning, allowing teams to co-optimize for both cost and quality . Caveat: this approach won't make sense in every domain, since general-purpose frontier models may be good enough . Article link: https://x.com/i/article/2090114065729503232

Great post on what post training looks like for applied AI use-cases to bring down costs and improve accuracy on certain tasks. This will… Update on our post-training effort
Sarah Tavel

Sarah Tavel (@sarahtavel, Benchmark partner) recommended @dittycheria's X thread on reference checks as "So much wisdom here on references," linking to it (https://x.com/dittycheria/status/2090414261545640095). Key frameworks in the thread: reference checks should attempt to disconfirm the case for the hire, not validate it ; prioritize blind references — former bosses, peers, direct reports, customers not on the candidate's list — over candidate-provided names ; discount praise by 3x and amplify criticism by 3x to counter softening ; and use forced ranking ("Top 1%? Top 10%?") to force real judgment .

So much wisdom here on references. Thank you [@dittycheria](https://x.com/dittycheria) for continuing to share yours! [https://x.com/ditt… Yesterday, the founder of a well-known, successful company called me to reference check two people he was considering for a senior role. …