The best daily AI content from around the web to get you caught up on developments before your first cup of coffee.

1 video, 20 articles

Executive Summary

Microsoft debuted its first streaming transcription model at No. 1 on Artificial Analysis, pairing real-time speech recognition with fast multilingual voice generation to support more responsive voice agents. Tavus is pushing the interface further with Griffin, a “human interaction model” designed to continuously interpret and generate speech, expressions, gestures, timing, and visual context rather than rely on turn-based exchanges.

AI governance tensions are rising alongside these advances. OpenAI reportedly fired researchers for allegedly sharing information with an external AI-safety group, highlighting the conflict between corporate confidentiality and independent scrutiny of frontier-model risks. California Governor Gavin Newsom also signed a law barring employers from relying solely on AI for hiring and firing decisions, creating a potential template for workplace-AI regulation elsewhere.

A new class of specialized decision models is emerging as an alternative to using general-purpose LLMs for every agent step. Clef, the open-source Kev project, and the 2-billion-parameter Strands Decider 2B directly score constrained options for tasks such as routing, classification, tool selection, guardrails, and risk assessment, promising faster, cheaper, deterministic, and more auditable systems. Related work on enterprise copilots proposes separating language-model recommendations from authorized execution, while AIM and AutoSynthData aim to make agent research and fine-tuning more systematic and verifiable.

Infrastructure and commercialization are advancing in parallel. Olmo-core 3 opens scalable training infrastructure for mixture-of-experts models reaching the trillion-parameter range, while reports of Google’s Gemini 4 Argon and faster OpenAI frontier models point to renewed competition on capability and latency. Shopify’s Canvas targets no-code bespoke storefront design, DoorDash is developing its own drone-delivery system to reduce costly human handoffs, and Albertsons is deploying OpenAI technology across internal operations and customer shopping—evidence that AI is moving from standalone assistants into core business workflows.

Trending Stories

Our first streaming transcription model debuts at no. 1 on Artificial Analysis | Microsoft AI

TLDR AIThe Rundown AI

  • Why it matters
  • Microsoft’s new speech stack combines top-ranked real-time transcription with fast, multilingual voice generation for more responsive voice agents.
  • Key details
  • MAI-Transcribe-2-Streaming supports 60 languages, produces partial transcripts in just over 100ms, and ranks No. 1 on Artificial Analysis for partial and final accuracy.
  • MAI-Voice-2.1 supports 23 languages and costs $22 per million characters; Flash generates 45 seconds of audio in 150ms for $15 per million characters.
  • Bottom line
  • Microsoft is competing on the full voice-agent loop—hearing and speaking—with a focus on accuracy, latency, multilingual consistency, and cost.

Exclusive | OpenAI Fires Researchers for Allegedly Sharing Information with AI Safety Group - WSJ

TLDR AIThe Rundown AI

  • Why it matters
  • The firings expose tension between OpenAI’s confidentiality controls and external scrutiny of its AI-safety practices.
  • Key details
  • OpenAI fired three safety-team researchers for alleged misconduct, according to people familiar with the matter.
  • The alleged violations included sharing confidential company information with a third-party AI-safety organization.
  • Bottom line
  • OpenAI is tightening control over sensitive safety information, even when disclosures involve outside safety groups.

Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs

TLDR AIHugging Face

  • Why it matters
  • Olmo-core 3 opens infrastructure that makes trillion-parameter mixture-of-experts training more efficient and accessible beyond major AI labs.
  • Key details
  • Scaling from 8 to 128 experts raised total capacity from 4.6B to 47B parameters while keeping ~3.2B active per token and cutting throughput by under 5%.
  • The new DDP-based stack delivered 2.7× higher throughput than its FSDP predecessor and benchmarked a 1.2T-parameter model across 512 NVIDIA B300 GPUs.
  • Bottom line
  • By combining expert and pipeline parallelism, distributed optimization, efficient routing, and MXFP8, Olmo-core 3 provides an open foundation for the next MoE-based Olmo.

YouTube

Latent Space

Recursive Language Models — Alex Zhang, MIT PhD

  • Why it’s interesting
  • Alex Zhang argues that today’s AI coding harnesses are largely variations on the same design—and that models could do much more if we redesigned the surrounding system rather than only scaling the model.
  • Recursive Language Models (RLMs) offer a surprising route to generalization: decompose large-context problems into reusable strategies and easier subproblems instead of forcing one model call to process everything.
  • Key concepts
  • Recursive Language Models: Systems that let a model programmatically inspect context, write code over it, and delegate simpler pieces to subagents, recursively reducing a difficult task.
  • Harnesses as computation architectures: Tools such as Claude Code, Codex, and Pi are opinionated programs that shape how a model searches, uses tools, and iterates—not merely interfaces around a fixed model.
  • Compositional generalization: RLMs can recognize that superficially different tasks, such as retrieval and aggregation, share the same underlying strategy, allowing training on one task to transfer to another or to longer contexts.
  • Alternative model designs: Systems such as looped transformers and fast classifier-style models challenge the assumption that every language model must be a standard autoregressive text-to-text decoder.
  • Main takeaways
  • Domain experts still provide major leverage: one knowledgeable person can steer a model away from reward hacks or wasted search that might otherwise consume billions of tokens.
  • Optimize systems end to end, not isolated components; the fastest individual GPU kernel may be worse overall if it prevents caching, fusion, or efficient coordination with later operations.
  • Academia’s comparative advantage is taking neglected, high-risk bets—not imitating frontier labs that possess vastly more compute and data.
  • Better harnesses should expose shared problem structure so models learn reusable procedures, rather than memorizing separate trajectories for every task.
  • Brute-force agent swarms can solve more problems, but cost and reliability remain central constraints; stronger decomposition and verification are often more valuable than simply spending more tokens.
  • Bottom line
  • The next major gains may come from redesigning how models recursively decompose, execute, and verify work—not just from building a larger autoregressive model.

No new videos: AI News & Strategy Daily | Nate B Jones, Lenny's Podcast, Every, No priors Podcast

Newsletter Articles

OpenAI cuts ties with 3 safety researchers, WSJ reports

via TLDR AI

  • Why it matters
  • The dismissals deepen scrutiny of OpenAI’s safety culture amid reports that executives sidelined internal security warnings.
  • Key details
  • OpenAI said it cut ties with three safety researchers after an investigation found they shared sensitive information outside approved procedures.
  • The researchers, outside organization, and information were not identified; it is unclear whether they first used internal reporting channels.
  • Bottom line
  • OpenAI frames the case as a confidentiality breach, but the timing raises fresh questions about how it handles safety-related dissent.

Our first streaming transcription model debuts at no. 1 on Artificial Analysis | Microsoft AI

via TLDR AI

  • Why it matters
  • Microsoft’s new speech stack combines top-ranked real-time transcription with fast, multilingual voice generation for more responsive voice agents.
  • Key details
  • MAI-Transcribe-2-Streaming supports 60 languages, produces partial transcripts in just over 100ms, and ranks No. 1 on Artificial Analysis for partial and final accuracy.
  • MAI-Voice-2.1 supports 23 languages and costs $22 per million characters; Flash generates 45 seconds of audio in 150ms for $15 per million characters.
  • Bottom line
  • Microsoft is competing on the full voice-agent loop—hearing and speaking—with a focus on accuracy, latency, multilingual consistency, and cost.

Introducing Clef: our open-source decision models, and new RL fine-tuning platform

via TLDR AI

  • Why it matters
  • Clef offers agents a faster, deterministic alternative to general-purpose LLMs for schema-bound decisions such as routing, classification, and risk scoring.
  • Key details
  • Cloudflare open-sourced Clef and Clef-flash under Apache 2.0, with vision support, a 64K context window, Jev API compatibility, and hosting on Workers AI.
  • Clef-flash posted 38.8 ms median latency versus Jev’s 524.1 ms across 43 evaluations; Cloudflare is also launching managed RL fine-tuning, with self-service tooling planned.
  • Bottom line
  • Cloudflare is positioning Clef as an open, low-latency decision layer for production agents, with domain-specific customization through reinforcement learning.

Introducing Kev

via TLDR AI

Why it matters

  • Kev offers open-source, self-hostable decision models that score answer options directly, enabling faster, auditable classification without token generation.

Key details

  • Four Apache-2.0 models span 0.8B–27B parameters; Kev-27B led the author’s six-panel evaluation, while Kev-4B and 9B approached it on several workflows.
  • Kev caches documents across questions; on an H200, 27B processes a 64k-token document in 9.4 seconds, then answers subsequent questions in about 0.7 seconds.

Bottom line

  • Use Kev-27B for maximum accuracy, but start with Kev-4B for local deployment or fine-tuning and validate thresholds on your own data.

perplexity-ai/pplx-decider-v1-27b · Hugging Face

via TLDR AI

  • Why it matters
  • Perplexity’s open 27B decision model turns prompts, text, and images into structured choices or yes/no probabilities with calibrated confidence.
  • Key details
  • Fine-tuned from Qwen3.8-27B, it averaged 85.71% across 11 benchmarks, beating Jev’s 84.51% and the base model’s 74.76%.
  • It led on FinancialPhraseBank, RAGTruth, TabFact, and Circa, but requires Python 3.12+ and roughly 49 GiB of GPU memory plus working space.
  • Bottom line
  • pplx-decider-v1-27b is a strong deployable classifier and routing model, though its gains vary substantially by benchmark.

AIM — Agentic Idea Management for Automated Research

via TLDR AI

  • Why it matters
  • AIM makes automated research more effective and auditable by explicitly organizing ideas, selecting experiments, and validating what was actually tested.
  • Key details
  • AIM combines an agentic surrogate, acquisition system, solution auditor, and resource planner to rank ideas, allocate parallel experiments, and reuse evidence.
  • Across 10 AutoLab tasks, AIM beat ScientistOne by 1.6 percentage points in system optimization and 4.9 points in model/CUDA development; on Flash Attention, it reached the baseline’s best score up to 3.1× faster.
  • Bottom line
  • Explicit idea management—not just stronger experiment-solving agents—can improve both research outcomes and the traceability of automated discovery.

Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs | Ai2

via TLDR AI

Why it matters

  • Ai2’s open-source Olmo-core 3 makes efficient trillion-parameter mixture-of-experts training more accessible beyond major AI labs.

Key details

  • A 47B-parameter MoE reached 52,000 tokens/sec/GPU on eight NVIDIA B300s—2.7× the prior stack—while scaling from 8 to 128 experts cut throughput by under 5%.
  • The system benchmarked a 1.2T-parameter model across 512 B300 GPUs at up to 858 TFLOP/s/GPU; MXFP8 boosted throughput 21% and reduced peak memory from 103 to 95 GiB.

Bottom line

  • By keeping experts GPU-resident and combining expert/pipeline parallelism, distributed optimization, and faster routing, Olmo-core 3 provides the open foundation for Ai2’s next MoE-based Olmo.

Introducing Strands Decider 2B: a small, open source, decision model

via TLDR AI

Why it matters

  • Strands Decider 2B offers a fast, local alternative to LLMs for routing, tool selection, guardrails, and other constrained agent decisions.

Key details

  • The open-source 2B-parameter model runs decisions in about 115 ms on an RTX 3090 and 153 ms for small tasks on an M3 MacBook.
  • Built from Qwen3.5-2B with a pointer head and LoRA adapter, it ranked 3rd of 33 in JevBench’s 2B class for accuracy and calibration.

Bottom line

  • Developers can use or retrain it locally to offload simple, repetitive agent decisions from slower, costlier generative models.

Griffin: The First Human Interaction Model | Tavus

via The Rundown AI

  • Why it matters
  • Griffin moves AI beyond turn-based voice assistants by interpreting and generating speech, expressions, gestures, timing, and visual context continuously.
  • Key details
  • In Tavus’s study, 48% of participants believed Griffin was human after a one-minute video call, versus a maximum 2% for its previous systems.
  • The full-duplex video-to-video model reassesses conversations at sub-second intervals, enabling interruptions, backchannels, emotional reactions, and real-time scene generation.
  • Bottom line
  • Griffin is an early but significant step toward AI that communicates face-to-face with human-like timing and nonverbal awareness; Griffin-Lite is in limited preview.

Search box to workflow copilot

via The Rundown AI

  • Why it matters
  • Enterprise copilots can move beyond answering questions to safely proposing authorized actions without letting language models directly control tools.
  • Key details
  • Algolia indexes and ranks task entities—manuals, workflow maps, validation forms, and tool triggers—using intent, permissions, risk, and environment constraints.
  • Retrieved deterministic tool contracts specify the exact tool, validated arguments, execution conditions, and expected outcome before a separate data plane acts.
  • Bottom line
  • Safe workflow copilots require governed, policy-ranked retrieval between natural-language intent and tool execution—not vector similarity or free-form model-generated calls alone.

OpenAI feels the frontier need for speed

via The Rundown AI

Why it matters

  • Frontier AI at near-real-time speeds could dramatically accelerate agent workflows, coding, and security investigations.

Key details

  • OpenAI’s invite-only Ultrafast API tier runs GPT-5.6 Sol at up to 750 tokens per second—14× its normal speed—using Cerebras compute.
  • Ultrafast completed Humanity’s Last Exam in 11 hours versus 78 hours for Fable with comparable results; pricing remains undisclosed.

Bottom line

  • If cost-effective at scale, Ultrafast could eliminate latency as a major constraint on frontier AI use.

Introducing Canvas: Opening the aperture on online store design

via The Rundown AI

  • Why it matters
  • Canvas could make bespoke Shopify storefronts accessible without coding skills, design expertise, or agency budgets.
  • Key details
  • The visual workspace shows every store page together and lets merchants use Sidekick to make real-time, store-wide code changes.
  • Sidekick validates code, reviews screenshots, and builds on experience from more than 25 million theme edits in the first half of 2026.
  • Bottom line
  • Canvas is rolling out now as an early-stage alternative—not yet a replacement—for Shopify’s existing theme editor.

Exclusive | OpenAI Fires Researchers for Allegedly Sharing Information with AI Safety Group - WSJ

via The Rundown AI

  • Why it matters
  • The firings expose tension between OpenAI’s confidentiality controls and external scrutiny of its AI-safety practices.
  • Key details
  • OpenAI fired three safety-team researchers for alleged misconduct, according to people familiar with the matter.
  • The alleged violations included sharing confidential company information with a third-party AI-safety organization.
  • Bottom line
  • OpenAI is tightening control over sensitive safety information, even when disclosures involve outside safety groups.

Our first streaming transcription model debuts at no. 1 on Artificial Analysis | Microsoft AI

via The Rundown AI

Why it matters

  • Microsoft’s new low-latency speech models could make multilingual voice agents faster, more natural, and cheaper to operate.

Key details

  • MAI-Transcribe-2-Streaming supports 60 languages, produces partial transcripts in just over 100ms, and ranks No. 1 for accuracy on Artificial Analysis.
  • MAI-Voice-2.1 supports 23 languages at $22 per 1M characters; Flash generates 45 seconds of audio in 150ms and costs $15 per 1M characters.

Bottom line

  • Microsoft now offers an integrated transcription-and-speech stack designed for real-time agents that can listen, act mid-sentence, and respond in a consistent multilingual voice.

Newsom signs law barring sole use of AI in hiring, firing decisions - UPI.com

via The Rundown AI

Why it matters

  • California is limiting employers’ use of AI in high-stakes personnel decisions, establishing a potential model for workplace regulation nationwide.

Key details

  • The No Robo Bosses Act bars employers from relying solely on automated systems to discipline or fire workers and requires human corroboration using additional information.
  • Employers must notify affected workers when AI was primarily used, disclose the employee data considered, and provide a human contact to discuss the decision.

Bottom line

  • AI may inform California employment decisions, but a human must review and substantiate disciplinary actions or terminations.

Argon aims to return Google to the frontier

via The Rundown AI

  • Why it matters
  • Gemini 4 Argon’s benchmark gains could return Google to the AI frontier after months without a top-tier flagship model.
  • Key details
  • Argon beat GPT-6 Astra and Claude Opus 5.5 on 13 of 19 Google-tested benchmarks and led DeepSWE coding with 77.9%.
  • Access is limited to vetted cybersecurity teams; API pricing starts at $2/$10 per million input/output tokens before doubling.
  • Bottom line
  • Argon looks competitive on paper, but Google’s comeback remains unproven until broader access validates its real-world performance.

DoorDash puts its own drone on the menu

via The Rundown AI

  • Why it matters
  • DoorDash is tackling the costly human handoffs that often undermine the economics of autonomous delivery.
  • Key details
  • DoorDash Air’s six-rotor drone uses a winch to collect and lower orders, with early deliveries averaging under five minutes.
  • About 80% of typical restaurant orders meet its size and weight limits, while software coordinates drones, couriers, and ground robots.
  • Bottom line
  • DoorDash’s success depends less on the aircraft than on making the entire kitchen-to-doorstep workflow cheaper and more reliable.

How Albertsons Companies is reimagining retail from the inside out

via OpenAI

Why it matters

  • Albertsons is deploying OpenAI across internal operations and customer shopping, testing how generative AI can improve retail decisions at national scale.

Key details

  • ChatGPT Enterprise and OpenAI APIs support employee workflows, product recommendations, promotional insights, commerce apps, and advertising across 2,200+ stores serving 36 million weekly customers.
  • A new Safeway experience lets shoppers turn recipes, photos, lists, or meal requests into product recommendations and carts in ChatGPT before checking out on Safeway.

Bottom line

  • Albertsons aims to use AI end to end—from faster internal decision-making to easier meal planning and grocery purchasing—with expansion planned across its other store brands.

AutoSynthData: Generating Training Data for Enterprise Agents

via Hugging Face

  • Why it matters
  • AutoSynthData turns enterprise-agent failures into validated, environment-specific training tasks, targeting weaknesses that generic datasets miss.
  • Key details
  • It generates task–system–verifier packages, then checks feasibility through execution, positive/negative verification, critique, repair, and batch-level diversity review.
  • Tasks are favored when the target succeeds in at most 1 of 3 trials while a stronger teacher succeeds in at least 2 of 3, keeping training near the model’s capability frontier.
  • Bottom line
  • ServiceNow’s pipeline creates an adaptive curriculum that shifts toward remaining weaknesses as the agent improves.

Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs

via Hugging Face

  • Why it matters
  • Olmo-core 3 opens infrastructure that makes trillion-parameter mixture-of-experts training more efficient and accessible beyond major AI labs.
  • Key details
  • Scaling from 8 to 128 experts raised total capacity from 4.6B to 47B parameters while keeping ~3.2B active per token and cutting throughput by under 5%.
  • The new DDP-based stack delivered 2.7× higher throughput than its FSDP predecessor and benchmarked a 1.2T-parameter model across 512 NVIDIA B300 GPUs.
  • Bottom line
  • By combining expert and pipeline parallelism, distributed optimization, efficient routing, and MXFP8, Olmo-core 3 provides an open foundation for the next MoE-based Olmo.