We can't find the internet
Attempting to reconnect
Something went wrong!
Hang in there while we get back on track
AI in EdTech Weekly
by avergin 92 sources
Weekly intelligence briefing on how artificial intelligence and technology are transforming education and learning - covering AI tutors, adaptive learning, online platforms, policy developments, and the researchers shaping how people learn.
AI in education is being forced to answer a harder question: not whether a model can produce a better answer, but whether the learner is still doing the thinking and can perform without it. Microsoft’s learning-mode rollout, a new synthesis of 274 studies, and emerging tutoring benchmarks all point in the same direction: the next product and procurement test is effort, transfer, and visible reasoning.
The product race is becoming a race over productive struggle
Microsoft says its Study and Learn Agent is available in Microsoft 365 Copilot for education licenses at no additional cost. Its stated design is deliberately different from a generic chatbot: it helps students understand, practise, study, and check their understanding using their own material, while preserving effort and productive struggle rather than doing the work for them. For K–12 students aged 13–17, an administrator must enable Copilot Chat; Microsoft says it plans to preselect Study and Learn for those tenants so Copilot Chat defaults to learning mode rather than answer-giving mode around the end of August.
The demonstration shows what that design means in practice. When a student is stuck on an equation, the agent asks what they have tried, simplifies the explanation, rebuilds toward the general method, and then gives another problem to test transfer. In writing, it asks what the student has already done and coaches the thesis process instead of drafting the thesis. It can also turn attached Word, PowerPoint, Excel, PDF, text, or OneNote material into activities such as flashcards.
Microsoft is also putting AI policy into the assignment interface. Teachers can set expectations per assignment—or as a default—with four suggested modes: no AI, brainstorming, editing only, or full AI, with editable text. That is a small but important shift from a general school policy to a task-level contract about what learning the assignment is meant to produce.
Renaissance is making a complementary, teacher-centered bet. Its Renaissance Intelligence product connects assessment data to core instructional materials and recommends adjustments for teachers; the company says about 400 districts were piloting or starting in July. It has deliberately deferred a student-facing AI model, arguing that AI tends to reach an answer too quickly and can undercut curiosity and productive struggle. The emerging split is clear: Microsoft is trying to make the student interaction more pedagogically disciplined, while Renaissance is using AI to help the teacher differentiate inside the shared classroom rather than sending students to isolated software.
The evidence is moving from “can it answer?” to “does learning survive?”
A podcast discussion of a new University of Sydney/Barker Institute synthesis reviewed 274 peer-reviewed studies published from 2022 through June on generative AI and school students. It reported a consistent distinction: AI can improve immediate performance in writing, mathematics, science, and coding, while producing better work with AI is not the same as learning better. In one highlighted mathematics study, 1,000 students with unrestricted AI access improved during practice but performed worse after AI was removed than students who had never used it. The review’s more useful question is therefore what thinking remains with the student: AI can help when it supports reasoning, explanation, and practice, but not when it simply supplies the answer.
The tutoring market’s own benchmarks point to the same constraint. EduClaw-Bench evaluated model-and-agent configurations over a 30-day learning horizon with simulated learners based on real knowledge-tracing data; almost none of the combinations tested, including GPT5.5 and Alibaba’s Qwen models, maintained strong tutoring performance over time. ELBench tested nine models across general capability, safety, basic education, and higher-order educational development, and the two education-specialized models in the study did not lead on the education-specific modules. The field still lacks a universally accepted definition of a good AI tutor.
That does not mean AI cannot perform difficult educational tasks. Ethan Mollick shared a field-study finding that an agentic loop using the now-obsolete o3-mini produced exam questions with psychometric properties on par with questions used on high-stakes standardized tests. The implication is not that AI-generated assessment is automatically trustworthy; it is that item-generation quality and student learning are separate evaluation problems.
The practical response is to make the learning process observable rather than rely on detection. A current assessment guide recommends collecting an initial question, claim, draft, peer-feedback reflection, revisions, and a final explanation; requiring students to explain choices; and using short conferences, audio reflections, debates, or gallery walks. It also recommends task-specific AI rules—such as allowing brainstorming but not drafting—and a short disclosure of how AI was used. The goal is not to make assignments artificially difficult, but to assess how students think as well as what they produce.
Students are becoming policy authors, not just policy subjects
At the America’s Youth AI Festival in Boston, student leaders from across the United States debated and passed the STUDENTS First Act, a proposed national K–12 framework that will be presented to AASA’s 10,000 member districts as a starting point for discussion—not as finished legislation. Their concerns included cheating, compromised critical thinking, privacy, security, mental health, and deepfakes, alongside equitable access and the benefits of self-directed learning.
The proposed boundaries are more nuanced than either blanket access or a blanket ban: no independent AI use before ninth grade; AI literacy beginning in grades K–5; no AI-generated written or artistic assignments; and older students allowed to use AI for brainstorming, studying, or editing with teacher permission and disclosure. The students also called for human review and an appeals process rather than sole reliance on AI-detection software. Their framework closely matches the direction of the strongest classroom guidance this week: define the purpose of the task first, then specify what AI use supports—or replaces—that purpose.
Workforce learning is becoming responsive to work itself
In frontline training, Opus describes a more operational form of personalization. The platform can ingest Google and Yelp reviews, analyze patterns by location, franchise, or individual, and automatically generate or retrain courses around the resulting performance gaps. Its Ask Opus agent answers from a company knowledge base and links the relevant training resource; in one large customer, 58% of responses had no resource attached, exposing missing knowledge—in that case, recipe data—and giving the company a way to identify what training to build next.
The experience is designed for varied frontline learners: Opus supports 130 languages and audio or reading modes, with 30% of users preferring audio-only training. A manager view creates a practical coaching checklist for employees who need support, coaching, or certification. The company’s CEO explicitly rejects engagement as the main measure, instead emphasizing knowledge retention and behavior change; claims such as reducing order-accuracy issues by 70% in 30 days are provider-reported and should be independently tested.
Andrew Ng’s new AI Engineering Skills Map points to the corresponding shift in what professional learning must teach. Based on more than 10,000 job postings plus expert interviews and surveys, it prioritizes building and deploying AI applications, software-engineering fundamentals, coding-agent fluency, and “shaping the build.” The map treats evaluation, error analysis, context management, tradeoffs, product judgment, and continuous learning as core skills—not optional technical polish.
What This Means
- For K–12 leaders: Make AI permissions task-specific and visible. A district should be able to say when students are expected to work unaided, when AI may support brainstorming or feedback, and what the teacher will see. Microsoft’s assignment controls and the student-authored framework show that “AI allowed” is too crude a policy.
- For assessment leaders: Require evidence of process, explanation, revision, and transfer—and preserve some opportunities for unaided performance. Immediate gains, polished outputs, and strong generated test items do not establish durable capability.
- For edtech buyers and investors: Treat an AI tutor as an intervention that needs longitudinal evidence, not as a chatbot with a friendly interface. Ask how it models learner progress, spaces practice, handles misconceptions, and measures what remains after the tool is removed.
- For L&D teams: Connect adaptive training to operational evidence, but measure retention and behavior rather than clicks. Multilingual and audio delivery can widen access; the unanswered-question log may be just as valuable as the generated lesson because it reveals where the organization’s knowledge base is weak.
Watch This Space
- Tutor quality standards: EduClaw-Bench and ELBench are early attempts to measure sustained tutoring, safety, and higher-order learning. Watch whether public evaluators and buyers converge on tests of retention, transfer, and learner modeling rather than one-turn helpfulness.
- From AI detection to provenance: The EU’s transparency rules now require labeling AI-generated images, audio, video, and certain public-interest writing; the accompanying discussion describes invisible text signals that can persist through copying and editing, while the generating company retains the decoding key. How schools handle provenance, false positives, and appeals will matter more than another detector leaderboard.
- Student-authored governance: The STUDENTS First Act is headed for discussion among AASA districts. Its uptake will test whether student voice can shape practical rules that protect authentic learning without treating all AI use as misconduct.
AI is moving into the routines that determine how people learn—lesson planning, differentiated materials, tutoring, assessment, and purchasing—but the support and quality controls around it are lagging. A spring survey found that 55% of teachers opposed classroom AI and 65% opposed students using AI for schoolwork, while a separate national poll found that 60% were already using it; 54% said AI makes critical thinking harder and nearly six in ten said it is eroding student–teacher trust. At the same time, 37% of surveyed K–12 decision-makers planned AI-tool spending for the next 12–24 months, alongside 68% planning educator professional development.
Trust is being lost in the guidance gap
The coexistence of use and skepticism is the important signal. Across six countries, teacher trust in AI was predicted by self-efficacy—not age or education level. Yet only 18% of U.S. K–12 teachers reported receiving formal, written AI guidance; 69% had no guidance for one-to-one tutoring and 58% had none for grading or feedback. Fewer than one in ten received formal guidance for any specific AI-assisted task.
Support changes the picture: 69% of teachers with official guidance said it encouraged AI use, compared with 51% who received only informal or verbal guidance. The same analysis recommends written policies, protected time for collaborative sensemaking, and objective assessment of AI literacy before training dollars are spent. The gap is global: more than 80% of surveyed Brazilian teachers reported only basic or limited AI knowledge, while more than 80% wanted continuous online professional development tailored to their level.
This makes the immediate implementation problem less “How do we persuade teachers to try AI?” than “What are they accountable for, what is safe, and how do they build judgment together?”
Assessment is shifting from the submitted artifact to the learning process
The Stanford Accelerator for Learning–ETS white paper Responsible Assessment in the AI Era frames learning as continuous, adaptive, and contextual, making single-point performance an increasingly incomplete basis for inference. It argues that AI could make formative assessment more feasible by synthesizing evidence and identifying patterns in large data streams—but also notes that LLM scoring may require more validity evidence and can be less reliable and more costly than traditional scoring.
Its practical design guidance is specific: capture revision, iteration, and collaboration; make AI feedback explainable; surface uncertainty; incorporate educator judgment; and avoid fully automated high-stakes decisions.
A Fairfield University implementation shows what that can look like. Biology professor Christine Rodriguez uses Blackboard’s AI Conversations to let nursing students practice explaining infections, tests, side effects, and home safety to an AI patient. The role-play gives students repeated, formative practice before later case-study assessments. Rodriguez reads the prompts and the students’ reflections, looking for whether they can synthesize knowledge and communicate it at a patient’s level.
Ethan Mollick makes the same distinction in a current interview: simply asking AI for help can remove the mental work, while a tutor should ask questions, explain gaps, and re-quiz without supplying the answer.
The emerging model is therefore not “AI grades the essay.” It is AI-supported rehearsal and evidence gathering, with the learner’s reasoning—and a human’s judgment—still visible.
The product and procurement race is now about context and quality
Back-to-school product updates show a common commercial bet: differentiated, standards-aligned lessons embedded in the teacher’s existing workflow. A current roundup reports that Claude for Teachers offers verified U.S. K–12 educators free premium Claude, standards-mapped curricula, nine tool integrations, and agentic workflows such as reviewing exit tickets and adjusting the next day’s lesson. Its stated data protections include no model training on teacher data and a FERPA-compliant data-processing addendum, but the program is currently for individual educators rather than district-wide rollouts.
Google Classroom’s updates add AI feedback, personal NotebookLM notebooks, standards tagging, a Classroom app inside Gemini, and forthcoming teacher-led guided-learning activities. Canva’s Learn Grid claims more than 50,000 curriculum-mapped resources, AI generation across 30-plus activity types and 16-plus languages, and multiple readiness versions of the same lesson. These features establish a workflow direction, not evidence of learning impact; the roundup itself describes a shared product bet rather than outcome studies.
Procurement is moving faster than policy. In the survey of 2,277 U.S. education decision-makers, the leading deal-breakers were subscription cost, state-standards misalignment, implementation complexity, and inability to integrate with existing systems. Fifty-seven percent reported no significant AI procurement-policy changes, even as leaders cited privacy, security, and the pace of technological change.
The quality problem is not hypothetical. Chalkbeat found obvious errors and flawed graphics in AI-looking Teachers Pay Teachers resources, a marketplace the parent company says is used by 85% of U.S. pre-K–12 educators. TPT’s parent company said algorithmic tools demote low-quality AI stores and that such products represent a “minuscule fraction” of sales; teachers and researchers nevertheless describe a market where time pressure and fluent-looking language make weak materials easy to miss.
One more grounded signal comes from outside generative AI. Carnegie Learning reports a matched comparison involving 144 Texas students in grades 6–8: after an average of 10 hours of live virtual tutoring, students gained 3.8 additional months of math learning, statistically significant, on top of regular instruction and the MATHia platform. It is a provider-reported virtual-tutoring result, not evidence that a generative model caused the gain—but it illustrates the kind of dosage, comparison, instructional design, and implementation detail that feature claims need.
Higher education is being asked to make AI rigor public
The U.S. Department of Education’s national call to action asks every postsecondary institution to publish, by the end of 2026, a public statement committing to reforms intended to restore trust. One of its seven guiding questions is how institutions will “incentivize rigor in the age of AI, combat grade inflation, and prioritize excellence in teaching and learning.”
A current governance proposal offers a concrete way to answer that question: separate assessment redesign from AI-literacy work. Its “Two-Lane Assessment Framework” distinguishes secured, supervised tasks for validating mastery from open, authentic tasks where AI use is assumed and integrated; suggested evidence includes oral defenses, process artifacts, reflections, and live applied performance. The proposal also treats AI literacy as curriculum—covering how systems work, their social and ethical effects, and discipline-specific co-reasoning—not merely a syllabus disclaimer.
What This Means
- For K–12 leaders: Publish task-specific guidance and fund collaborative teacher learning before expanding AI access. Screen-time decisions should begin with learning goals and student needs, not a universal clock; ISTE’s Richard Culatta argues that non-instructional screen use should be zero, while thoughtful use needs flexibility.
- For assessment leaders: Deliberately separate evidence of individual mastery from open-ended learning tasks. Use oral explanation, revision history, dialogue, and reflection where they reveal learning, and keep humans responsible for high-stakes decisions.
- For faculty and L&D teams: Make AI a structured tutor or rehearsal partner, not an answer engine. Require productive difficulty, feedback, reflection, and another attempt; Rodriguez’s nursing role-plays show how this can work in a discipline-specific setting.
- For product teams and investors: Standards alignment, privacy, interoperability, and low implementation burden are entry requirements. The differentiator must be credible evidence that learners retain or improve a capability, not simply that a system can generate more material.
Watch This Space
- Screen caps colliding with digital curricula: One Dallas–Fort Worth middle-school science teacher reports a 120-minute Chromebook cap while grades 6–8 science materials remain online-only, with no physical textbooks or lab workbooks available. A follow-up reports that administrators are tracking each student’s daily allocation, much of which is already consumed by reading and math software. This is a local anecdote, but it captures the implementation test for screen-light schooling: reduce low-value device use without removing the materials needed for instruction.
- Students as participants in AI governance: At Boston’s America’s Youth AI Festival, student leaders debated a proposed national K–12 AI policy. In the accompanying art competition, judges emphasized process, thoughtful AI use, and authentic student voice over impressive generated images; some students deliberately limited or avoided AI to keep parts of the work their own.
The central question this week is not whether AI can generate more educational content, but whether it can make learning more individual without making learners less capable. Andrew Ng’s LearnVector announcement and a set of classroom experiments point to the same design test: AI should guide a learning path, expose thinking, and keep a human accountable for the result.
Personalized guidance is the new product bet
Andrew Ng launched LearnVector with a $100 million investment from Coursera and plans to collaborate with Coursera and Udemy. He describes the company’s goal as moving from one-to-many education to a custom learning guide for each person—one that plans a path, adapts to how the learner learns, and stays with them until mastery.
The important part of the announcement is its warning as much as its promise. Ng says chatbots without guardrails can help students complete tasks and improve homework performance while leaving them less skilled through cognitive offloading; he also notes that chatbot advice is not always trustworthy. LearnVector’s proposed answer is not simply a more conversational chatbot, but a guided path built on authoritative learning materials.
A concrete school model points in the same direction. Montessorium in Texas says its students grew at 2.6 times the nationally expected MAP rate in year one, with language growth averaging about 3.5 times the national norm. The account is school-reported, so the result is better treated as an implementation signal than a general effect size. Its design is notable nonetheless: subject-matter guides teach alongside adaptive apps, computers are treated as one classroom material rather than the center of the room, and students have weekly individual conferences with educators.
The emerging product question is therefore not simply whether an AI tutor can personalize content. It is whether the system creates enough effortful practice, reliable content, and human relationship for personalization to become learning.
Assessment is shifting from polished output to visible reasoning
A Connecticut calculus experiment makes the shift tangible. Students had to explain the meaning of a derivative to “Devon,” an AI student with a specific misconception. Devon asked one question per turn, did not praise the student, and could return to an earlier confusion; the design prevented a learner from escaping with jargon or a definition that sounded right but could not survive questioning.
Twenty-one students ran the simulation 29 times. Scores averaged 82 out of 100, and the five students who voluntarily tried again improved by an average of 33.6 points. The beta also exposed its limits: one attempt was not enough, feedback and retry mechanisms were needed, and the AI scoring was not objective, requiring educator cross-checking.
At K–12 level, Vicki Davis describes a similar redesign. Her ninth graders use NotebookLM to produce podcasts from 20 sources, then explain the technology verbally without notes; she schedules three or four individual verbal interactions each semester. Nancy Frey recommends collecting the chatbot dialogue rather than only the final product, asking students to calculate the error rate of an AI-generated report, and using lateral reading to verify claims. For younger learners, she emphasizes “truth detective” and “information investigator” habits before direct chatbot use.
The practical response to AI-assisted work is becoming less about detecting a suspiciously polished essay and more about requiring explanation, source judgment, revision, and reflection. Timed writing and reflective questions such as what was most difficult also remain useful—not as a complete policy, but as ways to make the learner’s own thinking observable.
AI literacy is moving under the hood—and into institutional workflows
A sponsored Vicki Davis demonstration presents CodeAI, the nonprofit formerly known as Code.org, as a free tool for students to inspect how AI works rather than merely type into a black box. Its “glass box” curriculum deliberately creates moments worth examining, including culturally biased outputs and incorrect answers, then uses teacher-led discussion to analyze them.
CodeAI’s Karim Meghji argues that students should be in the driver’s seat: actively steering and questioning AI rather than passively remaining “in the loop.” The intended outcome is a healthy validation habit—asking whether to accept, challenge, and independently verify an answer. The product’s safety description is similarly candid: it uses pre- and post-filters and repeatedly tests hundreds of prompts, but acknowledges that no system will achieve complete protection against jailbreaks.
The same design logic is appearing in institutional tools. A Google for Education demonstration showed Gemini’s guided-learning mode asking questions instead of supplying a math answer; its Canvas tool generating HTML quizzes, interactive infographics, and apps that can be copied into Moodle or other HTML-capable spaces; and NotebookLM answering from selected sources rather than the open web.
Workspace Studio extends that shift from learning content to administration. No-code agents can be triggered by events such as an email, a file upload, or a calendar meeting, then use Gemini and NotebookLM to triage requests, consult institutional documents, and draft responses. The demonstration explicitly leaves human review at the end of the workflow.
Policy is beginning to catch up with these capabilities. Spain’s Council of Ministers approved a draft Organic Law on the good use and governance of AI, which still has to go through Congress. The accompanying higher-ed guidance recommends a centralized inventory of AI uses, human supervision, explainability, and impact and risk assessments before deployment.
Workforce learning is becoming an infrastructure problem
Andrew Ng argues that traditional education systems are too slow to prepare people for AI-shaped work: training should target the jobs of 2028 and beyond, including changing roles in marketing, recruiting, journalism, HR, and operations—not only software engineering.
That pressure is already visible inside universities. EDUCAUSE reports that AI use is moving faster than policy, training, and institutional strategy; employees are experimenting independently and may put proprietary information into unsanctioned tools. Its recommendation is to establish governance first, share ownership across IT, HR, and senior leadership, and make AI literacy an early practical step.
Access to advanced tools is also being widened at the research end. OpenAI says ChatGPT for Academic Researchers will give 10,000 scientists, mathematicians, and engineers free access to frontier models initially, expanding to 100,000 through 2027, with training and hands-on support. The program says researcher data will not be used to train models by default. That is an access initiative, not yet evidence that frontier-model access improves research or learning outcomes.
What This Means
For school and system leaders: Treat student-facing AI as an instructional-design decision, not a software rollout. A field report says education leaders are culling tools, prefer deeper implementation partnerships, and use a rough red/yellow/green logic: student-facing AI raises learning and safety concerns, teacher-facing AI needs alignment, and administrative AI is the easiest place to pursue efficiency. They also prefer pilots in their own context over vendor-produced research.
For assessment leaders: Require evidence of the process—explanation, source checking, dialogue, revision, and reflection—not only a finished artifact. The calculus beta suggests that retry and feedback are part of the learning design, while the K–12 examples show that transcripts and verbal defense can make understanding visible.
For edtech product teams and investors: Personalization and engagement are not impact measures by themselves. The field’s current warning is that AI may accelerate the old problem of confusing clicks and activity with learning; products need a credible account of what changes in learner capability, for whom, and under what implementation conditions.
For higher-ed and L&D teams: Pair access with governance and skill development. The useful unit of planning is not “which tool should everyone use?” but “which work is changing, what new judgment is required, and what protected workflow lets people practice it safely?”
Watch This Space
Independent evidence for AI-guided schools: LearnVector is a major commercial bet, while Montessorium offers an encouraging but school-reported result from one Texas campus. The next meaningful test is durable, independently evaluated learning—not another demonstration of personalization.
AI in high-stakes observation and exams: Anecdotal teacher reports describe proposed microphone-based evaluation, a 360-degree camera that rates teaching, and AI glasses colliding with testing rules and classroom-recording concerns. Responses range from district bans to legal consultation and instructions to act as if people are always being filmed. These reports do not establish prevalence, but they show that schools are having to define consent, accessibility, evidence, and accountability before the technology is settled.
AI moves from stand-alone chat to guided workflows
This week’s clearest development is the embedding of AI into the systems where teaching, support, and administration already happen. Google is rolling out Classroom as a connected Gemini app: teachers can ask natural-language questions about class activities, student performance, key concepts, and possible next steps. Gemini is designed to provide analysis and drafts—not to publish assignments or grade on a teacher’s behalf.
Google’s new Workspace Studio extends that shift to workflow automation across education editions, while allowing administrators to control access by organization, unit, or group and manage flow sharing. In one university demonstration, a flow identified helpdesk emails that might require mental-health support and routed those to a person or Chat space; less-sensitive requests received an AI-drafted response based on university information.
Microsoft is pursuing a similarly bounded model with its Study and Learn agent. Available to education users across licensing tiers, it is built around productive struggle, adaptive scaffolding, active learning, and application; it coaches rather than supplies answers, and can generate practice activities such as flashcards, quizzes, matching, and fill-in-the-blank exercises.
The important design choice is where automation stops. The strongest emerging pattern is not “AI runs the class,” but AI drafts, organizes, surfaces patterns, or supplies practice—while educators remain responsible for judgment, publication, assessment, and sensitive interventions.
Simulated practice shows promise when feedback is part of the learning design
A separate body of work points to a high-value use of AI: giving learners more chances to rehearse difficult interactions.
- In a 60-student medical-education study, students using virtual simulated patients improved their theoretical scores by about 6.5 points more than peers receiving traditional case-based training.
- In a study with 94 novice counselors, AI-patient practice plus AI feedback improved client-centered microskills; practice without feedback showed no improvement, and empathy declined relative to the feedback group.
- Columbia’s ACE negotiation system, tested with 374 participants, paired a bargaining counterpart with targeted coaching and outperformed both no-feedback and alternative-feedback conditions.
The shared lesson is that simulation alone is not the intervention. The useful sequence is attempt, feedback, reflection, and another attempt. A UC Santa Cruz history exercise using an LLM-enabled medieval-peasant role-play similarly found the activity worked best when embedded in discussion, analysis, reflection, and research—not treated as a one-off novelty.
This is a practical opportunity for professional learning, clinical education, teacher preparation, sales training, and language practice: AI can make rehearsal more available, but the feedback model and human-led debrief determine whether practice becomes learning.
The central learning question remains: does AI extend thinking or replace it?
Rose Luckin draws a useful distinction between beneficial cognitive offloading—where AI handles groundwork that enables more sophisticated thinking—and disadvantageous offloading, where learners bypass thought altogether. Her point is that assignment design determines the outcome: students use AI differently when they must demonstrate understanding rather than merely submit an output.
That distinction is visible in reading as well. An EdSurge commentary argues that AI text-leveling can deny learners access to complex texts and reduce the effort needed for retention and engagement. Its alternative is to keep the original text while using AI for supports such as vocabulary and syntax identification, phrase chunking, fluency-practice selection, and teacher think-aloud scripts.
Guided tools are increasingly built around this principle. Khanmigo, for example, is designed to refuse direct answers and instead use hints and reasoning checks; its math workflow compares a student’s proposed step with expected responses and asks the student to explain discrepancies.
“Learn fast, act more slowly.”
That is also becoming an implementation principle. In Linköping, Sweden, the municipality is preparing mandatory AI-literacy learning for 6,000 education staff before broader student rollout. Its framework emphasizes agency (human control), critical thinking, and responsibility for outputs and data. At the same time, Luckin warns that adoption and policy activity are increasing while gaps in access between advantaged and less-advantaged communities are also widening.
What This Means
For education leaders: Start with bounded workflows that solve a real problem—such as meeting follow-ups, source-grounded staff knowledge hubs, or draft communications. Name the person accountable for checking outputs before automating anything consequential.
For learning and L&D teams: Treat AI simulations as structured rehearsal. Build in a briefing, defined success criteria, feedback, reflection, and a repeat attempt. A conversational agent without those elements is not automatically a learning experience.
For instructional designers: Make student thinking observable. Ask learners to explain their reasoning, evaluate an AI response, revise it, or defend a decision. Avoid tasks where a polished output is the only evidence of learning.
For procurement teams: Examine more than headline capabilities. Luckin recommends asking what data trained a tool, what pedagogical premise shaped it, whether impact studies are well designed and independent, and whether the research population represents intended users.
For equity and accessibility work: Use AI to scaffold access—translation, vocabulary support, alternate formats, targeted practice—without quietly lowering the intellectual demand placed on some learners.
Watch This Space
AI-native school models: A $40,000-per-year private school opening in Oklahoma plans roughly two hours of AI-led core academics daily, with adult “guides” focused on projects and enrichment. It also says it cannot support students needing intensive behavioral, therapeutic, or one-to-one academic help—an early test of how access and inclusion play out in teacher-light models.
Classroom AI transparency: Google says educators will soon receive insights into how students interact with teacher-created Gems and NotebookLM resources, including topics and questions. The practical value will depend on whether those insights help teachers intervene rather than simply add another dashboard.
Automation in sensitive settings: Triage, student-support, and communications workflows can reduce routine workload, but they also raise the stakes for escalation rules, data governance, and human review.
Evidence beyond engagement: Simulation studies are encouraging, but education leaders will need independent evidence on sustained learning, equity, and implementation conditions—not only faster production of materials or positive user reactions.
Claude for Teachers raises the stakes—and governance questions
Anthropic launched Claude for Teachers, a free tool for verified U.S. K–12 educators. It offers standards-aligned lesson planning, personalized instructional materials, and analysis of student data; Anthropic says it incorporates standards from all 50 states. The company will pilot it in Detroit Public Schools Community District next school year, studying effects on educator well-being and practice. Students under 18 cannot access Claude directly.
The launch is significant less as a standalone chatbot than as another move toward institutionalized teacher AI. Claude joins offerings from Google, Microsoft, OpenAI, and Khan Academy aimed specifically at educators. Anthropic says verified teacher conversations will not train its models, and says it is working with the American Federation of Teachers on privacy and safety practices.
Detroit illustrates the operating model districts are beginning to define. Its new two-year teacher contract permits AI for instructional support such as lesson planning and drafting communications, but explicitly prohibits its use for grading or assessment decisions, IEP determinations, and discipline. The district must publish approved tools annually and provide ongoing professional learning on ethical use, privacy, bias, and inaccuracies.
That boundary is especially consequential as teacher use accelerates: an Education Week survey cited by Chalkbeat found 61% of teachers used AI in some capacity in 2025, up from 32% in 2024.
The practical divide: AI that guides learning versus AI that performs it
This week brought a consistent message across products, research summaries, and classroom accounts: learning outcomes depend on whether AI preserves productive effort.
Microsoft’s Copilot Study and Learn is designed around adaptive scaffolding, productive struggle, active learning, and application. In a hands-on review, it refused to become an answer machine or write essays, and it provided useful guided explanations for elementary particles and essay construction. But the reviewer also found that at least a quarter of supplied links failed and that the experience became less engaging over time than YouTube or Duolingo.
The evidence summarized by the AI in Education Podcast reinforces why this design distinction matters. A randomized trial with 1,700 secondary students in Sierra Leone reported a 0.26 standard-deviation improvement in maths for students using Gemini Guided Learning, with larger gains among those using it as recommended. In contrast, a study of 26,000 Chinese students found AI-assisted homework was completed 30% faster and earned 18% higher homework scores, but university entrance-exam results were 18–24% lower after two years.
A Brown University economics course offers a stark, though non-experimental, example of the assessment challenge: the class average fell from 96% on a take-home midterm to 48% on a proctored final; 18 students withdrew once the final was announced, and only four students scored 80% or above on the final compared with 57 on the midterm.
The emerging design principle is straightforward: use AI to question, prompt, explain, and give feedback—not to remove the learner’s responsibility for reasoning. A review of higher-ed research similarly found that stronger critical thinking predicts more positive AI-assisted learning outcomes, while students must retain responsibility for judgment.
“A co-pilot is only useful if trainees first learn how to be pilots.”
Sal Khan argues that this shift should elevate, rather than diminish, the educator’s role—from information delivery toward planning, coaching, motivation, and Socratic dialogue.
Tools are becoming more bounded, embedded, and evidence-oriented
The most actionable product developments are not general chatbots replacing a class. They are tools constrained by curriculum, source materials, or structured workflows.
- Microsoft Teach now combines curriculum planning, modification of existing content, assessments, and interactive learning activities. Microsoft says it has exceeded 2 million users since its October 2024 release and supports standards in 62 countries. Its Copilot Notebooks can ground chat and study materials in files that users upload; associated study guides can generate quizzes, flashcards, and fill-in-the-blank activities.
- Microsoft Learning Zone is now available on all Windows 11 PCs, creating interactive lessons with self-paced activities, feedback, and teacher insights. It supports up to 10 slides per lesson on standard Windows 11 machines and up to 99 on Copilot+ PCs.
- At Faculdade de Medicina da USP in Brazil, faculty loaded NotebookLM with teaching materials, question-writing guidance, and validated past questions to create residency-exam simulations. Professors reviewed the outputs; the school reports reducing production of a 120-question exam from at least 240 hours to under 10 hours, and expanding from two to eight simulations.
These cases show a useful pattern: source-grounding, specialist workflows, and human review can make AI more usable than open-ended prompting. They do not eliminate the need to inspect outputs. TeacherServer, for example, provides more than 1,194 pre-prompted education tools and says inputs are deleted immediately, but a review found its generated slides basic and its writing edits sometimes inaccurate.
Access and inclusion depend on implementation details
For bilingual learners, chatbots can keep core content in English while providing explanations and scaffolding in a student’s home language. Educators interviewed by EdSurge also described using chatbots for 24/7 practice, targeted sentence starters after initial attempts, text annotation, and family conversations through simultaneous translation.
But these benefits are conditional. AI speech systems may be less accurate with accented speech, dialectal variation, or language-switching, which can reduce support quality or produce inaccurate assessment. Home-device and internet access remain a second risk of widening gaps.
The same requirement for human context appears in student support. AI can identify patterns and reduce administrative work, but advisors—not algorithms—determine how support should be delivered.
What This Means
- For district and school leaders: Treat AI adoption as a governance and professional-learning program, not a tool rollout. Detroit’s contract offers a concrete starting point: define approved purposes, reserve consequential decisions for people, publish the tool list, and train staff continuously.
- For teaching and learning teams: Prefer tools that create productive struggle—questioning, feedback, practice, and revision—over answer generation. The difference can separate short-term task performance from retained learning.
- For assessment leaders: Build more opportunities for students to explain and defend their work. Caltech admissions is reportedly using AI to flag cases for oral examination of research-paper understanding, while evidence on detectors suggests they should be treated as a signal for review, not proof of misconduct.
- For product teams and investors: The strongest near-term implementations are bounded: aligned to standards, grounded in uploaded materials, or built for a narrow instructional workflow. The differentiator is increasingly the surrounding privacy, review, and deployment model—not just model capability.
- For workforce and lifelong learning providers: Programs are moving toward practical AI fluency and durable skills. Khan Academy’s planned Constellation Institute proposes group simulations and peer review around communication, collaboration, creativity, critical thinking, and leadership.
Watch This Space
- Detroit’s Claude pilot: Its teacher training and planned study of well-being and practice could provide a meaningful test of a teacher-only AI model operating alongside formal labor protections.
- Proactive tutoring: Khan Academy is developing a more proactive Khanmigo after finding that some students did not know how to ask the first version for help.
- Assessment that makes thinking visible: Oral exams, process evidence, and conversational simulations are gaining attention. “Friction Bots,” for example, use deliberately difficult branching conversations to assess argumentation and have been tested across 12 institutions.
- Teacher-facing AI at scale: As systems add planning, differentiation, and assessment support, the critical questions will be whether staff receive the training they need and whether implementation protects teacher judgment. Microsoft’s survey found a substantial gap between the training leaders believe is available and what teachers and students report experiencing.
The evidence is sharpening: AI should prompt, listen, and coach—not simply answer
This week’s most consequential signal is practical: AI’s learning value depends heavily on the job it is asked to do. In a Beijing randomized trial with 148 kindergarteners, a chatbot using structured dialogic questions matched a trained human reader on story comprehension and word learning; a chatbot that only narrated the story produced the lowest comprehension scores.
The pattern reverses when AI does the cognitive work for students. In an English schools trial of 344 students aged 14–15, students who could ask an LLM questions about history texts performed worse three days later on retention, comprehension, and free recall than peers who took notes—even though 90% used the tool to request elaboration. Penn research cited by Ethan Mollick similarly found that unrestricted AI use left students believing they had learned when they had not, while purpose-built AI tutors produced large learning gains.
“You don’t learn very well when people just give you answers.”
For reading, the clearest near-term fit may be listening and guided practice. Carnegie Mellon’s Reading Tutor outperformed sustained silent reading in a seven-month comparison across 178 students in grades 1–4. More recently, a Texas study of 107 students with dyslexia found that consistent use of BuddyBooks was associated with nearly double state-test growth; the study was correlational, not causal.
Generative content remains less dependable. Expert reviewers judged many GPT-4o reading questions suitable for operational use, but only 42.6% targeted the requested inference skill; models also could not reliably simplify passages to a fourth-grade level. The practical distinction is increasingly clear: structured questioning and fluency support have stronger evidence than open-ended answer bots or unreviewed content generation.
Two policy paths: Illinois sets a framework; New York City pauses purchasing
Illinois issued a non-mandatory, 400-page framework for responsible K–12 AI use. It places human interaction at the center, defines AI as a tool to inform rather than replace teaching, calls for civic engagement with communities, and asks districts to set deliberate, locally determined purposes. The guidance includes grade-level examples for lesson planning and prompt engineering, alongside treatment of privacy, transparency, cultural bias, and hallucinations. District policy templates, professional learning, and internet-safety instruction related to AI-generated cyberbullying are planned over the coming school year.
New York City has taken a more restrictive interim step. Chancellor Kamar Samuels asked principals to halt new educational-software purchases until the Education Department finalizes revised AI guidance later this summer. The freeze follows criticism of the city’s initial guidance and calls for an AI moratorium from parents, educators, and more than half of City Council members. Officials have also struggled to identify which AI-enabled products schools already use because many purchases happen at school level; a survey is now underway.
The pause may complicate summer planning because schools use software for core functions as well as academic support. Together, Illinois and NYC illustrate two immediate governance needs: define acceptable instructional uses, and know what technology is already in classrooms.
Assessment is shifting from AI detection to visible thinking
The response to AI-generated work is increasingly about designing assessments that require students to demonstrate their own reasoning. One proposed “Assessment Puzzle Framework” layers text, visuals, annotations, voice reflections, and personal connections; each layer makes a pasted AI response less sufficient as evidence of learning.
That approach aligns with Mollick’s recommended division of labor: use AI tutors for work outside class, then protect in-person time for discussion, active learning, essays, and role-play assessments. Instructors using Wikipedia assignments are making a related shift: rather than focusing solely on whether a tool was used, they check whether students’ claims and citations are verifiable.
There is urgency behind the redesign. One veteran teacher reported seeing AI complete 40% of homework, while also reporting that students in an AP computer-science course where AI was deliberately integrated have all passed the AP exam over three years. Her stated priority is “learning detection” rather than AI detection. That is an anecdotal account, but it reinforces the research distinction: the educational result depends on whether AI replaces practice or is incorporated into a structured learning process.
AI’s strongest institutional role may be making human attention more targeted
Equal Opportunity Schools combines student survey data—including belonging, trusted adults, and aspirations—with academic records and AI-driven predictive analytics to identify students overlooked for advanced coursework. A Mathematica study found partner schools identified more than 2.5 times as many underrepresented students ready for AP, IB, or dual enrollment as traditional methods, with identified students performing as well as peers once placed.
That is a useful model for “human-led, technology-augmented” practice: AI can surface patterns, while educators decide how to respond. It also offers a contrast to systems that treat personalization as automation. Students may benefit from tailored information, but institutions still need to make them feel understood rather than managed by an algorithm.
For teacher workflows, NotebookLM is being positioned as a bounded alternative to open-web prompting: it works from documents users upload and can turn source materials into audio overviews, explainer videos, slide decks, mind maps, flashcards, quizzes, and study guides. Its constraint is also its value: outputs are limited to the materials supplied, so it is best suited to transforming and exploring a known set of sources rather than replacing source selection or educator review.
AI use is also becoming a student-wellbeing issue
Nearly two-thirds of teens report experimenting with AI, and some use chatbots for companionship, romantic relationships, or unvetted mental-health support—not only schoolwork. Educators are being urged to ask students how and why they use chatbots, identify needs those interactions may be filling, and connect students with healthier human support where appropriate.
This expands AI literacy beyond prompts and plagiarism. It includes explaining pattern recognition, helping students verify outputs with human judgment, and asking them to show both an AI output and the revision or reasoning they contributed afterward.
What This Means
- For classroom design: Favor tools and prompts that require explanation, retrieval, revision, and repeated practice. Avoid treating fluent answers as evidence of understanding.
- For district leaders: Pair AI guidance with procurement visibility, staff learning, and explicit expectations for privacy, safety, and instructional value. Illinois’ framework and NYC’s purchasing freeze show different routes to the same operational problem.
- For assessment teams: Build process into the assignment—oral explanation, annotated sources, draft history, personal connections, or in-class performance—rather than relying on a detector after submission.
- For equity initiatives: Use data systems to widen access to opportunity, but keep relationships and educator judgment central to the intervention.
- For families and student-support staff: Treat chatbot relationships as a topic for candid conversation and AI literacy, not just a screen-time or academic-integrity issue.
Watch This Space
- NYC’s final AI rules: The promised revised guidance—and how it handles younger students and existing software—will be an important test of whether districts can move from broad principles to workable controls.
- Reading-fluency evidence: AI listening tools have a long evidence lineage and promising implementation data, but the field still lacks a modern causal study to establish their current impact.
- Rural implementation capacity: A Texas Tech study found AI professional development remains scarce in rural schools. Action-oriented programs that start with teacher-identified classroom problems are emerging as one response.
- Tutor modes in mainstream tools: ChatGPT’s study mode is now activated by typing “@ study,” while Gemini also offers a study mode. The practical question is whether these modes consistently preserve learner effort as they reach more students.
Human agency is becoming the organizing principle
The week’s clearest shift was practical: more of the strongest education voices are defining good AI use around what stays with the learner and educator, not just what the tool can automate . Rose Luckin argued that AI’s biggest contribution to education is the questions it forces institutions to ask about human agency, accountability, and the limits of automation . Her framework is direct: preserve human agency so AI never makes critical decisions alone, be transparent about what systems cannot do, and put institutional accountability in place before deployment; she noted that the EU AI Act classifies education AI as high-risk .
“The most important thing that AI has done for education is not about what the AI can do. It’s about the questions that it’s forcing us to ask.”
Luckin also made two implementation points that showed up elsewhere this week: don’t mandate adoption ahead of staff capability, and don’t confuse performance gains with learning gains . She pointed to South Korea’s rollback of an AI textbook rollout after educators said they lacked sufficient training, and stressed that even a well-designed tool depends on rollout context and staff capability .
That same logic is now shaping classroom design. Ethan Mollick again pointed to a recurring pattern: doing the homework matters, AI tutoring that supports classwork can help, but AI “help” that reduces mental effort harms learning . Brooklyn teacher Rayhan Ahmed described the classroom version of that risk as students using AI to bypass “the intellectual and emotional struggle of learning” . And Alpha School’s MacKenzie Price offered a design principle rather than a product pitch: the screen should be used only for the narrow window where it can calibrate difficulty and catch misconceptions, while “the cognitive work must remain with the child” .
One practical higher-ed response came from Wikipedia assignments. Faculty interviewed by Lance Eaton described these projects as a way to move students from AI “black box” consumers to transparent knowledge producers, focusing them on neutrality, sourcing, edit histories, and public accountability . Wiki Education’s current boundary is also instructive: students should not use AI to draft Wikipedia content because AI-generated text overwhelmingly fails verifiability checks, even when citations look plausible . AI can still help earlier in the process by surfacing content gaps or hard-to-find sources .
Policy is getting more specific because student use is already mainstream
The policy story this week was driven by scale. Across multiple surveys, AI use is no longer marginal in education . Common Sense Media found that 86% of teens have used generative AI and nearly a quarter use it daily, with schoolwork among the top uses . In higher education, Lumina and Gallup found 87% of students use AI, 57% use it daily or weekly, and 42% of students at colleges that discourage AI still use it regularly . An Australian survey of 10,000 university students found daily use rising to 20%, weekly use now a majority, and a shift toward writing improvement and content summarization rather than simple Q&A .
What has not kept pace is guidance. Oxford University Press reported that only 40% of students think using AI for all homework counts as cheating, while 75% want teachers to use AI more and only 15% say they have received enough school guidance . The AI in Education Podcast summarized the broader pattern bluntly: students are using AI at high rates and remain confused about the rules .
Regulators are starting to answer that gap with more specific policies. Norway is imposing severe restrictions on AI use in primary schools, allowing only limited supervised use in lower secondary and broader access from age 17 . In Australia, the higher-education regulator updated its guidance on AI in assessment, recommending version-history evidence or process-tracking tools such as Cadmus, Inktrail, and Turnitin Clarity—and explicitly warning that AI detectors alone are not enough for misconduct allegations .
Taken together, these moves suggest a more mature policy phase: less abstract debate about whether AI belongs in education, and more operational decisions about when students should use it, how institutions can inspect the process, and where human judgment must remain decisive .
The strongest product signal from ISTE: AI is moving into teacher workflows, coaching, and controlled classroom use
The most interesting tool pattern this week was not fully autonomous teaching. It was AI embedded into narrower, better-defined educational jobs . Tech & Learning’s ISTE 2026 winners were notable less for broad claims than for “thoughtful” integration into teaching and learning . In K-12, that meant age-appropriate creative tools such as Adobe Aqua for elementary learners , AI literacy sequences that run from foundational skills through ethics and algorithmic reasoning in Learning.com’s EasyTech platform , and hands-on introductions to coding and AI through LEGO Education kits and scaffolded lessons .
The teacher-workflow layer is getting more concrete too. MagicSchool and MagicStudent were highlighted for fitting daily staff workflows while supporting student independence rather than simply giving answers . Microsoft Learning Zone was cited for combining lesson creation, live delivery, personalization, and formative insights while keeping educators in control . Learning Genie’s Curriculum Genie emphasized standards alignment, UDL support, differentiation, and a teacher-approval model . On the hardware side, BenQ and Samsung both emphasized on-device or built-in AI in classroom displays rather than separate add-ons .
The governance layer of the market is growing just as quickly. Brisk Teaching stressed student and teacher data privacy . Securly focuses on monitoring AI usage and sentiment to help districts enforce policy . In higher ed, Airia centers visibility, security, compliance, and auditability for institution-wide AI use , while D2L Lumi positions itself as a controlled AI extension inside Brightspace .
A separate teacher-facing development came from the AI2S project. The system analyzes classroom video to classify math questioning patterns, cognitive demand, and student engagement with accuracy said to rival trained human observers . Its key design choice is that it is not positioned as a dashboard. It guides teachers through a coaching cycle to identify a focus area, review lesson data in context, and build an action plan .
“We’re not trying to build teachers a Fitbit. We’re trying to build them a coach.”
Privacy is central to the pitch: teacher video and feedback stay private, and classroom data are not used to train the models . Pilots are underway in Texas, New York, and Virginia, with expansion planned into reading and language arts .
The limitation across all of these tools is the same: they show where the market is heading, but not yet definitive proof of learning impact. Many are expert-selected award winners, workflow tools, or pilots rather than independently validated outcome studies .
Lifelong learning tools are accelerating, but the human layer still matters
Outside formal schooling, AI is clearly compressing the time required to build and deliver learning experiences . Duolingo CEO Luis von Ahn said two employees with no chess or engineering background used AI to build the first version of Duolingo’s chess course in about six months, and that the course now has 7 million daily active users . He also described an internal rule that AI should be used only when it benefits learners, with productivity gains allowing the company to produce more educational content .
At Duolingo, AI is also changing how teams work: product managers now use AI to prototype ideas instead of relying only on written documents, the company runs AI training days, and it stopped trying to score AI use in performance reviews because that encouraged superficial behavior instead of better outcomes .
Von Ahn’s counterbalance is worth noting for anyone treating this as a teacher-replacement story. He argued that AI can be excellent at repetition and adaptation, but teachers remain better at motivation, context, and inspiration .
New study formats are also emerging for self-directed learners. NotebookLM introduced Short Video Overviews that turn source materials into 60-second vertical explainer videos, first for paid subscribers and then for all web users in English . Coursera’s new Ollie app, available to Coursera Plus subscribers, packages short lessons with interactive practice, conversational AI support, and hands-free listening or exploration modes . Even individual tool-building is becoming more accessible: EdSurge profiled LibraryAid, a personalized book-recommendation app built via vibe coding by David Webb, who had no prior computer-science background .
What This Means
- For K-12 leaders: Age matters more now. Norway’s staged restrictions, Luckin’s emphasis on human agency, and product designs like Adobe Aqua and EasyTech all point toward more age-specific AI rules rather than a single district-wide posture .
- For higher ed: Assessment policy is moving away from detector-led enforcement and toward process evidence. Version history, public-facing assignments, oral defense, and transparent knowledge-production tasks look increasingly practical .
- For teachers and school systems: The strongest near-term AI use cases are not “teach the class for me.” They are workflow support, coaching, personalization with educator control, and privacy-aware governance tools .
- For curriculum and workforce planners: AI literacy is being framed less as prompt tricks and more as engaging with, creating with, shaping, and managing AI—while keeping human agency and values visible . In parallel, the World Economic Forum’s readiness work says “willingness to learn” is now the top transversal skill in nearly 20% of European job ads, and education ranks as the second-most AI-intensive sector .
- For product builders and investors: The market signal is toward narrower, auditable, domain-specific products. Tools that can show clear workflow value, privacy protections, and human oversight are gaining more traction than general-purpose classroom chat alone .
Watch This Space
- Process-aware assessment tools: Australia’s regulator has now put institutional weight behind version histories and workflow evidence. Expect more assessment products to make process visible by default .
- AI coaching for teachers: AI2S is still in pilot mode, but its expansion from math toward reading and language arts is worth tracking if it can scale feedback without losing trust .
- Short-form study media: NotebookLM’s video overviews and Coursera’s Ollie both point to a new layer of AI-generated study companions built around quick explanations and interactive follow-up .
- AI-first content creation in lifelong learning: Duolingo’s chess course is a concrete example of AI compressing course-development cycles. Similar patterns are likely to surface across microlearning and workforce training .
- AI-native upskilling programs: Gauntlet AI’s 10-week training model, and its use of tools like NotebookLM, Obsidian, and reusable “skills” playbooks, suggest that workforce learning may increasingly teach people how to research, orchestrate, and document with AI—not just how to use a chat window .
Grounded AI is becoming the product default
The biggest shift this week was structural, not model-related: new education tools are increasingly being tied to course materials, district resources, and specific workflows instead of handed over as open chatbots. Gemini’s new study notebooks are grounded in class materials, start with a diagnostic quiz, build bite-sized interactive lessons, update based on follow-up quiz results, and roll out globally on the web at no cost. They also sync sources and chats with NotebookLM.
NotebookLM, meanwhile, added fully customizable flashcards so learners can edit questions and answers, add new cards, and share sets with classmates.
At the institution level, the pattern is even clearer. MagicSchool says a Georgia state audit found 58% of teachers have used its product, and the company has now added district-level controls to ground AI in local resources. Brisk presented a similar idea through "Curriculum Intelligence" built from district guidance and curriculum libraries, while TrekAI positions itself as a supervised "learner’s permit" and Lightspeed gives districts visibility into what AI tools students are already using.
Higher ed is testing the same move with different infrastructure. Denison built a token-priced multi-model environment with 17 models for students, faculty, and staff, and requires incoming students to complete a foundational AI course focused on ethics and use.
These launches are promising, but they are still mostly product rollouts and operational examples rather than independent proof of better learning. Access is also uneven: Gemini says mobile and school-issued accounts are coming later this summer, and district or campus systems depend on local setup, policy, and budget.
The learning design principle is getting clearer: don’t remove the struggle
This week’s strongest research signal was practical: AI helps most when it guides thinking instead of collapsing it. In Ghana, the Rori tutor produced an effect size of 0.36 at about $5 per student by giving hints and guidance before solutions. A Carnegie Mellon study found answer-giving chatbot use tracked with worse exam performance, while answer-withholding proof review tracked with better outcomes. In a GPT-4 field experiment with nearly 1,000 high-school math students, the answer-giving version led to 17% worse later exam scores without AI than the control group, while a safeguarded version avoided that drop. A UK RCT from Google, LearnLM, and Eedi found AI support increased novel problem-solving by 5.5 percentage points over human tutors alone.
"Ask me one question at a time, waiting for my answer in between, to help me think through this problem, to help me discover angles I haven’t thought of."
That advice from Mindstone’s Joshua mirrors the research trend: use AI as a thinking coach, not a search replacement or answer machine.
The limitation matters just as much. One physics-feedback system was wrong in about one case in five, and even very strong students often failed to spot those errors. The broader evidence still supports a narrower claim than the market sometimes implies: AI reliably improves performance while learners have access to it, but durable unassisted retention and transfer remain unsettled. Rich scaffolding also helps weaker learners more than stronger ones, which means the same design will not fit every student.
Alpha is trying to turn AI-native schooling into a scalable schedule
New Alpha School interviews pushed the conversation beyond tutoring and toward full schedule redesign. The model uses AI to assess knowledge gaps, keep students working at roughly 80-85% difficulty, and compress core academics into a two-hour daily sprint with no homework. Alpha says that model has produced 2.6x faster learning, top 1-2% performance, and an average SAT of 1535 for 11th graders.
The more distinctive claim is about motivation. Alpha’s founders argue that motivation is "90% of the solution," and say "Time Back" lets students spend afternoons on workshops, sports, entrepreneurship, and life skills. They report that 96% of students say they love school, while guides focus on one-to-one motivation and support rather than lectures and grading.
The new scaling detail is access. Alpha says AI costs are falling from about $10,000 per student toward the hundreds, and it is using Texas vouchers to expand through sports academies and gifted programs while planning inner-city public-school pilots in Texas. Its longer-term goal is a tablet-based product priced below $1,000 that can teach a full curriculum in under two hours a day. Those public-school pilots are still upcoming, with the founder saying the data will be published when they open.
Trust, bias, and policy are moving to the center
The policy story this week was not faster adoption but slower, more contested adoption. New York City delayed final AI guidance until later this summer after its March draft drew nearly 6,500 comments and broader backlash. The draft used a traffic-light framework that banned AI for assessments and grading while allowing lower-risk uses such as brainstorming lesson plans, but it said little about student use. Officials are now considering age-based expectations as they try to prepare older students for an AI-present world without letting AI replace their thinking.
Bias concerns are getting harder to dismiss as abstract. Victoria Hedlund reported that AI explained a simple science concept with more scaffolding but less technical rigor to girls than boys, awarded lower marks to "Victoria Hedlund" than "Victor Hedlund" on the same GCSE history paper, and gave girls more discouraging physics career advice than boys with the same qualifications. She also argued that names, locations, and inferred socioeconomic context can lower the level of challenge an AI tutor offers, and that the frequency of AI interaction can expose students to bias far more often than a human teacher would.
The operational response is becoming clearer: minimize identifying inputs, avoid using AI for marking or detectors without strong evaluation, and treat outputs as results to inspect rather than conclusions to trust. That lines up with broader guidance from Tech & Learning, where Microsoft’s Matt Jubelirer argued that AI literacy now goes beyond prompting to judging capabilities, context, and ethical use, and that grading still requires human judgment.
"We don’t want humans in the loop. We want humans in the lead."
That same principle is showing up in product architecture. e-Literate argues that multi-agent AI can fit academic work because it mirrors teams of specialists, but each extra agent adds token costs and lossy context handoffs, making large-scale deployment harder to budget and audit.
The AI literacy debate is widening into a human-skills debate
A second strategic shift this week was philosophical. Justin Reich argued that education still lacks evidence on what effective AI literacy practice actually is, and that many early frameworks repeat the same mistake schools made with web literacy: packaging a lofty new skill bundle before expert practices are clear. In his view, domain expertise may matter more than generic knowledge about how large language models work, so schools should run local experiments and compare evidence of learning rather than assume a ready-made AI literacy playbook exists.
At the same time, industry voices are shifting from "teach AI" to "teach the human capabilities AI makes more valuable." Executives speaking to educators emphasized collaboration, resilience, communication, negotiation, leadership, critical inquiry, and ethical reasoning as future-proof skills, even as employers and coalitions like RAISE US push AI-enabled training and workforce transition support.
For lifelong learners, Andrew Ng’s advice was notably old-fashioned in the best way: start with efficient coursework, then build small projects, take handwritten notes to improve retention, and make learning a regular habit rather than a burst activity.
What This Means
- For K-12 systems: the near-term winners are likely to be grounded tools with clear guardrails, visibility, and age-appropriate rules—not open-ended chat alone. District leaders have more reason to prefer tools tied to curriculum, local resources, and data-minimization standards.
- For higher ed: AI literacy is looking less like a single standards document and more like a sequence of local experiments, explicit policies, and assessment redesign. The bar for using AI in grading or high-stakes judgment should stay high.
- For teachers and learning designers: this week’s evidence again favored AI as a critic, coach, or draft partner over AI as an answer source. Study notebooks, guided tutoring, and structured planning are moving faster than fully automated teaching.
- For learners and workforce teams: structured AI tools can make practice easier to start, but they do not remove the need for error-checking, domain knowledge, and regular study habits.
- For investors and product builders: demand is moving toward grounded, workflow-specific AI, but bias evaluation, cost control, and proof of learning impact are becoming as important as model quality.
Watch This Space
- Grounded consumer study stacks: Gemini’s study notebooks, NotebookLM’s editable flashcards, and free AI tools from Khan Academy, CK-12, and Pear Start all point to a more structured self-study layer emerging on top of general-purpose models.
- Public-school and lower-cost AI-native models: Alpha’s planned Texas public-school pilots and its tablet-based mass-market ambitions are worth tracking if they move from founder claims to published data.
- Bias audits and age-based governance: NYC’s delayed rewrite and Hedlund’s experiments suggest next-wave policy will focus less on blanket approval or bans and more on age, data minimization, and bias exposure.
- Agentic systems with humans in the lead: multi-agent course design and support tools are advancing, but whether institutions can afford them at scale remains open.
- Human-skill-first workforce learning: the tension between AI fluency and durable human skills is likely to shape both curriculum design and employer training over the next cycle.
The big development: education is drawing a harder line between answer machines and learning tools
This week’s clearest signal was not a new model. It was a sharper distinction between AI that helps learners think and AI that helps them avoid thinking .
Ethan Mollick pointed to a recurring pattern: students naturally reach for AI on homework, but off-the-shelf chatbots act like assistants, not tutors, by providing answers that reduce mental effort and undermine learning . He also cited a large study in China showing that when AI shortened homework time by lowering effort, test scores fell too .
“Across studies, a theme: AI tutoring in support of classes is good, using AI to ‘help’ with homework is bad.”
The same tension is showing up in writing and assessment. One higher-ed analysis argued that AI has widened the college-readiness gap in writing by making it easier for students to produce polished text without doing the thinking writing is supposed to develop . A narrower use on the Packback platform looked more promising: AI handled grammar and style feedback so instructors could focus on ideas, and student writing improved modestly over a semester . EdSurge’s podcast reached the same practical test for K-12: whether students are learning to think with AI or using it to bypass productive struggle .
Sam Altman said he expected schools to redesign quickly after ChatGPT, with projects that require AI but still stretch thinking, yet he still sees no significant systemic change across education 3.5 years later and warned that, without redesign, critical thinking skills could atrophy .
Mollick’s practical response is more concrete: more in-class assignments, AI tutors that challenge rather than answer, and prompts that use AI as a critic during debate or argument rather than as a completion engine .
Even commentary on higher ed outcomes is moving in this direction. The AI in Education Podcast highlighted Berkeley research suggesting that A grades rose 30% in take-home writing- and coding-heavy courses after ChatGPT, while one computer science course’s failure rate reportedly rose from 7% to 35% when students later faced exams without AI . Researchers from Australia, New Zealand, and China are now explicitly arguing for Socratic AI companions that prompt reflection and understanding, rather than vanilla chatbots that act as a crutch .
The response is shifting from bans to explicit learning design
Higher ed and K-12 are starting to turn that insight into design choices instead of generic rules. Lance Eaton argues that year 5 of generative AI in higher education should be the year of program-level curriculum change, because students are still graduating after four years of scattershot experiences: one professor requires AI, another bans it, another treats it as misconduct, another never names it . His recommendation is not to simply teach the tool or ban it, but to map where students should first encounter AI, where they should use it, where they should work without it, and where they should learn refusal, verification, disclosure, and judgment .
A related UK discussion is pushing even further upstream. Rose Luckin argues that if AI can master knowledge-heavy curricula faster and more accurately than humans, schools need to shift toward richer human capabilities: creativity, problem-solving, metacognition, resilience, empathy, and better assessment of those skills .
K-12 policy is also getting more operational. Microsoft Teams now lets educators set assignment-level AI expectations: full AI use, editing only, brainstorming only, or no AI use, with customizable labels and defaults . Students see the guideline when they open the assignment and, if their school enables it, a direct button to open Copilot .
That product choice matches district-level policy thinking. Tech & Learning highlighted a three-part framework for districts starting from scratch: understand how students and teachers are already using AI, protect student data and personally identifiable information, and address academic integrity without taking students out of the driver’s seat of their own learning . It also draws a distinction between prohibition-heavy acceptable use policies and responsible use policies that explain reasoning and treat students as participants . In Massachusetts, Shrewsbury Public Schools built five pillars around student preparation, student learning tools, staff tools, guardrails, and academic integrity, aligned to the district’s Portrait of a Graduate rather than to any single product .
This is also showing up in professional development. Hillsborough County Public Schools, a district serving more than 200,000 students, put 1,000 educators through a summer week of training on responsible use of MagicSchool AI . The underlying message is similar to Monica Burns’ advice: start AI guidance with the kind of thinking you want students to do, not compliance alone .
Tools are getting more specific — and their limits are clearer
The most credible product activity this week was not “AI does everything.” It was AI being inserted into narrower learning workflows .
In Microsoft’s education stack, Learning Zone lets educators attach AI-built interactive lessons directly to Teams assignments, render them inside Teams for students, and provide built-in checks and feedback . Those lessons can also draw from partner content including NASA, Figma, and Minecraft . Rubric generation is becoming more constrained as well: when teachers create a rubric with AI, the standards already attached to the assignment are automatically carried into the rubric . The limitation matters: AI lesson generation requires a Copilot Plus PC, even though students can complete the lesson in Teams on their side .
At the school operations level, the near-term gains remain mostly administrative. One principal described using AI daily to turn state memos into slide decks, teacher texts into parent messages, long emotional emails into summaries and replies, contract PDFs into queryable answers, and scattered event details into calendar entries — saving “a few hours” a day and freeing more time for students and staff . Monica Burns describes the same division of labor more generally: AI can do the drafting, formatting, and structuring, but human review, edits, and knowledge of students still shape the final product .
For self-directed learning, platforms are getting more interactive. Copilot Notebooks is now available without paid student Copilot licenses, and one suggested use is uploading a curriculum to generate study guides, activities, and infographics . Google NotebookLM is already being used by students to turn class slides into self-quizzes and answer keys, reinforcing retrieval practice . Andrew Ng is pushing toward a more conversational model in CodeDream.ai, where learners interact through simulated video calls and embedded JavaScript demos instead of passively watching static videos . But Ng’s verdict is restrained: online learning tools are better than they were 10 years ago, not yet truly transformed .
Adoption, meanwhile, remains a constraint. Mollick says AI interfaces like chatbots, Codex, and NotebookLM are not intuitive in practice and contain “a dozen little tricks and traps” that block effective use . He also says many people never get past the difficult first hour, which keeps AI in the “kind of like Google” box . Chalkbeat’s reporting on a Stanford AI tutor study shows what that looks like in schools: human guidance increased use by only 1-4 minutes a week, many students never logged on, total time stayed far below the 30 minutes a week needed for reading gains, and there was no meaningful difference in reading scores .
“The challenge isn’t just building good AI tools. It’s really getting students to use them, and that seems to take the same type of intentional design that we’ve learned matters with other ed tech interventions and tutoring.”
AI-native models are expanding — but not all in the same direction
Some schools are no longer treating AI as an add-on. They are designing schedules, staffing, and pedagogy around it from the start.
At Alpha School, students spend two morning hours on personalized academic work with AI tutors or adaptive apps that give immediate feedback and let students advance on mastery . Afternoons shift to four hours of team-based life skills, projects, collaboration, and conversation . Alpha argues that this split lets AI handle individualized cognitive work while freeing more time for authentic social learning, and says its classes rank in the top 1-2% nationally . Reporting from Michael Horn’s microschool series reinforces the broader pattern: the most interesting schools are not simply maximizing AI use; they are being explicit about where AI belongs and what human capabilities — autonomy, entrepreneurship, deep research, feedback, and strong foundations — they still want school to build .
Alpha is also trying to rebut a common critique directly. Its leaders say AI-first schooling should produce more thinking, not less, and are pairing the model with explicit humanities work, including students reading Tocqueville’s Democracy in America and debating it for the age of AI as part of building “philosopher-builders” .
Higher education is seeing its own AI-native experiments. The AI in Education Podcast highlighted a new Italian online university built from scratch around AI optimization, serving 112,000 students with 400 academic staff . That is a radically different staffing model from legacy higher ed. But Andrew Ng’s comments are a useful counterweight: what people need to learn is changing quickly — coding agents, AI building blocks, and broader product skills — yet the delivery of training is still being reinvented in real time .
Economics are starting to shape product design
For edtech buyers and investors, one of the most practical notes this week came from e-Literate: current AI economics do not fit education’s usual software model .
Schools and colleges budget around fixed annual costs, while metered AI introduces variable usage that can spike unpredictably . e-Literate points to enterprise examples of blown token budgets, revoked licenses, and large unexpected bills as signs of what happens when usage limits are weak . The likely consequence is product design, not just procurement, changing: mainstream platforms are more likely to ship constrained AI actions, narrow buttons, predefined workflows, and usage caps than open-ended magic text boxes or expensive multi-agent systems . That slowdown may frustrate some vendors, but it could also reduce the odds that education scales the wrong tools too quickly .
What This Means
- For K-12 leaders: Redesigning homework and assessment is now harder to avoid. The evidence and commentary this week point toward more in-class checks, explicit AI-use expectations on assignments, and tutor-like AI that preserves effort instead of replacing it .
- For higher ed: Course-by-course AI rules are too inconsistent. Program-level maps of where students should use, refuse, verify, and disclose AI are becoming a more practical governance model .
- For teachers and L&D teams: The best short-term use cases remain bounded ones — interactive lesson building, standards-aligned rubrics, administrative drafting, and study supports — with human review still central .
- For edtech builders and investors: Capability is not enough. Products need engagement, usability, and cost discipline. Low-usage tutoring pilots and unsustainable token economics can kill otherwise promising ideas .
- For learners: AI is most useful when it behaves more like a critic, coach, or quiz-maker than an answer machine .
Watch This Space
- Socratic companions and public-interest tutors: Researchers are pushing companion-style AI that promotes reflection, while Mollick argues universal tutors are now technically plausible if built with public R&D, transparency, and the right scaffolding .
- Mainstream platforms embedding guardrails: Assignment-level AI labels in Teams suggest more classroom software will make AI expectations visible inside the workflow, not just in policy documents .
- Next-generation study platforms: Khan Academy says its next launch will combine trusted content with AI tools to help students persist through hard learning, while Copilot Notebooks and CodeDream point to more interactive self-study formats .
- AI-native school builders: Alpha’s summer internship and new engineering cohort show schools investing directly in building learning apps, not just buying them .
- Cost-shaped AI design: Expect more predefined AI actions, fewer unlimited chat interfaces, and closer scrutiny of whether usage actually translates into learning gains .
The clearest signal this week: structured AI is outperforming generic AI use
In Sierra Leone, Google DeepMind positioned AI as a response to teacher shortages, describing it as a partner that can extend educators’ reach without replacing them . Over eight weeks, students increasingly used Gemini to understand concepts rather than just get answers, with problem-solving queries rising from 68% to 90% . EdSurge also pointed to a Sierra Leone study in which a one-day AI training for secondary teachers was followed by math gains equivalent to more than a year of additional schooling .
"AI can act as a partner to support educators in these environments – amplifying their reach without replacing their essential expertise and skills."
The contrast with generic chatbot use is getting sharper. Standard free LLMs can lower brain activity and retained learning by encouraging what one report called "cognitive surrender," while a carefully designed AI tutor in an undergraduate physics course produced twice the learning gains of active, in-person instruction . Estonia’s emerging policy follows the same logic: students build foundational knowledge first, then use AI later in the learning process for feedback and assisted learning; earlier grades are deliberately excluded for now .
Schools are building tighter AI workflows instead of relying on public chatbots
In Australia, some schools are already doing this themselves. In Broken Bay Diocese, a Year 6 student built a science agent inside a controlled "kids’ pool" environment; it checks a learner’s level, adapts explanations or tests, and can even add engagement cues like jokes. The class later adopted the agent because it worked across different learning needs .
Another school built a secure Gemini + Apps Script tool that combines GPA, testing, and timetable data so teachers can query class or student breakdowns and get differentiation suggestions without moving student data into public systems . At Cathedral College in Rockhampton, a lesson-starter agent was grounded in school teaching frameworks, curriculum documents, and teacher-contributed examples, but its creator emphasized that cultural preparation came first and that the tool was not meant to replace human mentoring .
"It can't exist on its own. It's not meant to be a standalone agent or replace a human coach or mentor."
District tutoring programs are moving in the same direction. Newark Public Schools received $400,000 from New Jersey to expand high-impact tutoring that uses AI with teacher oversight for math and reading . District leaders said they expanded Khanmigo after pilot users showed math-score improvement .
The next wave of tools is more lesson-native than chatbot-native
Microsoft’s new Learning Zone turns educator prompts, uploaded files, or vetted resources such as OpenStax into interactive lessons in minutes . The lessons combine bite-sized content slides with multiple exercise types, immediate feedback, retries, and conditional "nested" slides that give students extra practice when they miss a concept . Teachers can use it for topic introductions, wrap-ups, flipped learning, or live instruction with anonymous aggregated knowledge checks, then assign lessons through codes, links, Teams, or an LTI-compatible LMS and review performance reports afterward .
The capability is notable, but the limits matter too. Students can access lessons in a browser on any device, while lesson generation currently requires a Copilot+ PC, with a broader trial planned . Microsoft also says more generation languages and an in-class teaching mode are coming .
Assessment is still where AI most clearly hits its limits
The week’s most useful assessment research separates scoring from feedback. In a randomized trial across 178 schools in Brazil, AI essay scoring performed at the level of human review; students improved by about a tenth of a standard deviation whether essays were scored by AI alone or AI plus human graders . But feedback is a different task: it requires identifying what a specific student got wrong in the context of that student’s reasoning .
That gap shows up across multiple studies. In middle-school math, the model that scored best produced teacher-preferred feedback only 12% of the time, while GPT-4 produced the most trusted feedback and the worst scores . On science work, LLMs matched teachers on next-step "Feed Forward" guidance but lagged on "Feed Back" that diagnoses the student’s specific reasoning error, scoring 3.05 versus teachers’ 3.52 out of 5 . In a physics-feedback study, about 20% of AI responses were inaccurate, and students rated wrong feedback as just as accurate as correct feedback .
This helps explain why AI is reallocating teacher time rather than eliminating the teacher role. In Brazil, AI scoring gave teachers about 30% more one-on-one writing conferences without increasing workload . And it helps explain Justin Reich’s warning that AI is most useful when the user already has enough domain knowledge to separate strong output from confident nonsense .
Real classrooms are already adapting. Reich’s 120-interview project found widespread homework bypass, with some students deciding which assignments are important enough to do themselves and teachers responding with everything from rewrites and detectors to harder AI-required tasks . Teachers on Reddit describe moving essays and tests in-class or on-demand because "anything that goes home is AI’d now" . Another teacher working with younger students reported near-identical AI-generated responses on a research assignment .
Higher ed and policy are moving from experimentation to governance
In higher education, roughly three-quarters of faculty say students use GenAI to write essays and papers, and roughly the same share of faculty use GenAI themselves . But most institutions are still in a wait-and-see phase or running patchwork experiments, not seeing broad gains in learning or efficiency . That is pushing redesign in two directions: tougher assessments such as live oral defenses, presentations, and more rigorous feedback loops , and better student-support systems that connect academic, financial, and well-being data while reducing administrative overhead . It also aligns with the argument that higher education and workforce systems need continuous reskilling, more personalized learning, and more real-world experience as AI changes work .
Governance is becoming more explicit. EDUCAUSE described low-risk AI uses such as brainstorming, tutoring, translation, summarization, and simple coding; medium-risk uses such as course design, grading, student feedback, and administrative tasks; and high-risk uses involving student records, HR data, or financial information . For community colleges, the warning is that commercial tools may encode assumptions that do not fit part-time, working, or caregiving students, which can automate weak judgments at scale . Recommended safeguards include contextual auditing, vendor transparency, and collaborative governance with faculty and student advocates .
The same shift is happening at system level. England’s Department for Education has funded 16 edtech firms to build trustworthy AI tools for lesson planning and marking using a national curriculum data/content store prototype . A £1 million pilot has now expanded into a £23 million, four-year program recruiting 1,000 schools and colleges as test beds for AI and edtech . Google DeepMind and Edy have announced a randomized trial of Learn LM with 1,500 students across 10 schools in England , while OpenAI is testing a learning-outcomes measurement suite with 20,000 students in Estonia and Anthropic and CodePath are running a 15-month classroom study across thousands of students .
Scale, though, is not the same as settled evidence. Ben Williamson argues that these programs can privilege signals of impact and scalability over broader forms of evidence, while reshaping classrooms into measurable testing sites for experimental AI products . That warning matters because AI economics are not like normal software: every use carries inference costs, private or local deployments add storage, cybersecurity, hardware, networking, and technical expertise, and districts still have few examples of what universal access would actually cost . UNESCO’s Digital Transformation Collaborative is one example of the response, framing digital transformation around coordination, connectivity, cost, capacity, content, and data .
What This Means
- For school systems: Separate AI for tutoring and practice, AI for lesson creation, and AI for assessment. The same tool may score reliably and still fail at diagnostic feedback .
- For instructional leaders: Training and workflow design matter more than access alone. Sierra Leone’s gains followed teacher preparation, and Australian implementations started with secure environments and cultural work—not just tool rollout .
- For higher ed and L&D teams: Access without redesign leads to patchwork. The durable opportunities are stronger assessment, better student support, continuous reskilling, and more real-world AI use through projects, internships, co-ops, and apprenticeships .
- For buyers and investors: Ask what must be true locally—curriculum fit, teacher prep, infrastructure, language, inclusion, data protection, and affordability—before treating pilot results as scalable .
- For learners and families: In this week’s coverage, AI literacy was defined less as basic tool use and more as questioning outputs and evaluating reliability; foundational knowledge still determines whether AI helps or misleads .
Watch This Space
- Teacher-made microtools: "Vibe coding" is lowering the barrier for teachers to build task lists, translation workflows, dashboards, and practice games. One fourth-grade teacher reported an AI-built review game that led to students scoring five points higher on average with no retests .
- More frontier-lab trials inside education systems: The DfE test-bed expansion, Learn LM trial, OpenAI’s Estonian measurement suite, and Anthropic’s CodePath study will shape both evidence and market expectations .
- AI-native school models: Alpha World School’s launch points to a more ambitious version of AI-enabled schooling, pairing daily AI-driven academics with fieldwork in Kenya and Ecuador and research projects with university faculty .
- Agentic tools for self-directed learning: NotebookLM’s new research companion can build a source repository from loose questions, surface its reasoning process, and export outputs in formats from charts to spreadsheets and documents .
- Cost and governance as adoption bottlenecks: As pilots broaden, recurring inference costs, privacy demands, and local hosting decisions may matter as much as the model itself .