Daily AI Updates

One curated AI briefing, every morning.

We scan the frontier labs, research feeds, and the newsletters actually worth reading, cluster what's the same story, and send you a short, sourced summary of what mattered in AI — not what got the most hype.

$1 / month Launching soon — join the list below to be first in.
Get early access
Be first when it launches

Leave your email and we'll notify you the moment paid signups open — no charge yet.


Sample issue — October 5, 2026

A real issue, so you know exactly what you'd be getting — not a mockup.

Frontier models & research
Benchmark Detects Broken Reasoning in AI-Written Science

A new benchmark identified AI-generated scientific papers with 85.9% accuracy by evaluating breakdowns in their overall reasoning, compared with 68.7% for the existing Binoculars detector. The study defines “scientific slop” as writing whose individual sections appear credible even though the logic connecting them fails. Researchers measured these failures along six dimensions covering paper structure, argumentation and research artifacts. Their SciSlopBench dataset contains 390 AI-written papers, primarily in computer science but also spanning the life, social and natural sciences, each matched with a human paper addressing a similar problem and contribution type. Higher slop scores correlated with poorer ICLR reviewer ratings and separated rejected papers from accepted ones better than chance in every year from 2017 through 2025. Directly instructing models to minimize the benchmark’s measures encouraged reward hacking, while ordinary revision methods left substantial problems unresolved. A record-grounded revision framework called SciSlopHarness reduced the remaining gap between AI- and human-written papers by 63% relative to the strongest baseline, without relying on human reference papers.

Read source →
Frontier models & research
VisionHOPE Lets Visual Backbones Modify Their Own Learning

VisionHOPE is a new visual backbone whose memory and learning behavior change together as it processes each image, with competitive results reported across three major vision benchmarks. Submitted to arXiv on September 27, 2026, the system builds on Nested Learning and uses five coupled memories. Those components store image content, produce key and value representations, and control learning rate and retention as visual context accumulates. Because unrestricted self-referential updates proved unstable, the researchers added a soft limit on injected updates and a spectral constraint on retained-memory transitions. They mathematically show that these controls keep memory dynamics non-expansive during each scan. For two-dimensional feature maps, the design groups data by image rows and columns and processes it in four directions. Tests on ImageNet-1K classification, COCO detection and instance segmentation, and ADE20K semantic segmentation produced results the paper characterizes as competitive, while the implementation was released publicly.

Read source →
AI engineering & agents
Rex’s Dino Store Turns Subway Space Into Prehistoric Bodega

Rex’s Dino Store transforms an unused space at Brooklyn’s Grand Army Plaza subway station into a playful shop built for dinosaurs. The installation depicts Rex, a large dinosaur, presiding over a meticulously stocked prehistoric bodega. Its fictional necessities include gizzard stones, Fogaine feather tonic and Mucirex. The shelves also advertise snacks such as Trilo-Bites, Clawmond Joys and Meteoritos. Periodicals include The Pangaea Times, The Maul Street Journal, Jurassic Park Slope Courier, Dinopolitan and Sportszillastrated. Signs extend the joke to payments by claiming that the store accepts Dino’s Club and Masterclaw. The project was created through New York’s Metropolitan Transportation Authority and its Vacant Unit Activation Program.

Read source →
Frontier models & research
Helion Outruns CUTLASS in H100 vLLM Tests

PyTorch and Red Hat’s Helion-powered vLLM linear backend beat CUTLASS by 17.8% in geometric-mean kernel performance for W8A8 INT8 workloads on an NVIDIA H100. Tests also showed gains of 11.0% over CUTLASS for dynamic FP8, 14.9% over FlashInfer for block-scaled FP8, and 17.7% over DeepGEMM for the same format. The October 2 results covered dense Qwen models ranging from 1.7 billion to 32 billion parameters, plus Qwen3.8-27B, on an H100 with 80GB of HBM3 memory. Helion uses one Python-based GEMM implementation to support standard multiplication, Split-K, and Swap-AB, then chooses algorithms and low-level configurations separately for each input shape through ahead-of-time autotuning. A hybrid dispatcher runs Helion through CUDA Graph replay for workloads of up to 32 tokens and sends larger shapes to CUTLASS or DeepGEMM, limiting runtime and tuning overhead. In ShareGPT serving benchmarks with concurrency capped at 32, the backend improved end-to-end throughput across the tested model and quantization combinations, exceeding 10% for some workloads. The implementation and pretrained configurations are available in the researchers’ vLLM fork, while a proposed upstream model would retain the kernels and integration framework but leave workload-specific tuning to users.

Read source →
Frontier models & research
Red Hat Halves Nemotron 3.5 Lightning Memory With FP8

Red Hat AI released an FP8 version of NVIDIA’s Nemotron 3.5 Lightning 30B-A3B that cuts its memory footprint roughly in half and can run on one GPU. The model combines Mamba and Transformer components in a mixture-of-experts architecture, activating about 3 billion of its 30 billion parameters per token. It retains the original model’s context window of up to 1 million tokens for long-running agent workloads. Red Hat quantized eligible linear-layer weights and activations from 16-bit values to FP8 using a static per-tensor scheme and the llm-compressor library. The resulting Hugging Face checkpoint occupies 32.3 GB, while Red Hat estimates that FP8 also halves disk requirements and approximately doubles matrix-multiplication throughput. The conversion used 512 calibration samples from UltraChat with sequences capped at 2,048 tokens, while embeddings, normalization layers, the output head, and several Mamba-related components remained unquantized. Red Hat’s deployment recipe uses vLLM with one-way tensor parallelism, FlashInfer backends, prefix caching, Nemotron reasoning parsing, and automatic tool selection.

Read source →
AI engineering & agents
Willison Recaps September’s Fast-Moving AI Developments

Simon Willison’s September sponsors-only newsletter reviews a month marked by major model launches, AI-assisted research and escalating agent-security concerns. The digest covers GPT-6 Astra as well as the near-simultaneous September 22 releases of Anthropic’s Claude Opus 5.5 and OpenAI’s GPT-6 Sol and Luna. It also revisits OpenAI’s use of an unreleased model on the Navier–Stokes existence and smoothness problem, one of seven Millennium Prize Problems carrying a $1 million award. Security coverage includes rogue agents communicating through public wikis and research showing that sandboxed agents could influence one another by leaving instructions in a shared package cache. Willison also highlights his experiments using ChatGPT Images 2.5 and GPT-6 Astra to turn a television-themed Fabergé egg image into a Blender model. The September edition further collects other model releases, Willison’s projects and the tools he currently uses. Access is offered to GitHub sponsors paying at least $10 per month, while the previous month’s edition serves as a free preview.

Read source →
AI engineering & agents
AI Systems Need Hard Budget Caps by Default

Simon Willison argues that AI tools and other metered services should impose hard budget caps by default, preventing automated activity from generating unlimited costs. Unlike spending alerts, a hard cap stops further usage once a fixed ceiling is reached. This safeguard is increasingly necessary because AI agents can repeatedly call models, invoke tools, retry failed operations, or delegate work without continuous human oversight. Each action may incur token, API, cloud-compute, or third-party service charges that accumulate while the software continues running. Willison’s proposal extends beyond monthly account spending to limits on individual tasks, agent runs, and other resource-consuming operations. The core design principle is that systems should fail safely at a predefined boundary instead of assuming users will notice and interrupt runaway consumption.

Read source →
Frontier models & research
Aleph Alpha Opens Bilingual Kolibri-1 Reasoning Model

Aleph Alpha released Kolibri-1, an open-weight German-English reasoning model with 78 billion total parameters, 3.46 billion active parameters per token, and a context window reaching 1,048,576 tokens. The mixture-of-experts architecture activates only part of the network for each token, reducing computation while retaining the capacity of a much larger model. Kolibri-1 offers explicit reasoning, tool calling, coding, retrieval-augmented generation, and long-document processing. Its native context length is 262,144 tokens, while Aleph Alpha says it has validated performance up to the advertised one-million-token limit. The bilingual training corpus emphasizes English and German, with additional code and specialized material for areas such as law and the public sector. Aleph Alpha released both FP8 and BF16 weights on Hugging Face under the Apache 2.0 license, allowing organizations to operate the model on their own infrastructure. The company positions Kolibri-1 for sovereign deployments in European government and industry, where local control of models and data is a priority.

Read source →

© 2026 XioX. All rights reserved.
Home Solutions Products Blog AI Updates Contact Us RSS