Library / Research

Research

49 explainers · newest first

Research explainer
🎧№ 1206
RESEARCH
Source · Bain & Company

Why most CEOs' AI efforts are just 'fake progress': The seven decisions that determine real enterprise AI transformation

Most companies run plenty of AI pilots, but they never change an end-to-end business process, and they never turn a one-off win into a capability they can reuse next time. The dividing line is whether the company ends up with its own data, its own ways of working, and its own record of what it has learned.

Enterprise AIAI agents▶ Play08-23 · 18 min read
Research analysis
№ 1203
RESEARCH
Source · Simon Willison / Promptwatch⚑ Research

ChatGPT Search’s underlying retrieval just changed: ChatGPT is now doing large-scale, targeted site scraping that upends GEO strategy

A single answer is no longer just a broad sweep of the web. ChatGPT may now discover candidate sources first, then lock onto specific sites to dig deeper. As the retrieval order shifts, GEO expands from page-level ranking into a three-layer contest: domain inclusion, page matching, and citation conversion.

ChatGPT SearchGEO08-22 · 11 min read
Free
Research analysis
№ 1196
RESEARCH
Source · Reuters⚑ Research

Personalized mRNA cancer vaccines pass a major milestone: AI helps design them, and the era of one drug per patient is getting closer

Moderna and Merck's personalized mRNA cancer vaccine hit positive results in a Phase 3 trial for melanoma. Algorithms help select each patient's tumor neoantigens, then a custom vaccine is manufactured for that individual. This is the first time a truly personalized approach has cleared a large-scale clinical hurdle, but widespread use still faces three big challenges: data, cross-cancer efficacy, and cost.

AI drug discoverycancer vaccines08-20 · 8 min read
Free
Research explainer
№ 1195
RESEARCH
Source · Cursor

How Cursor rebuilt Git hosting to tackle a 20-year-old problem: Inside the Continuity architecture

From Git's DAG and packfiles to GitHub Spokes, and then to Continuity with S3 WAL as the source of truth: a full walkthrough of why large-scale Git hosting is hard and how Cursor solves it.

CursorGit08-20 · 18 min read
Free
Research Analysis
🎧№ 1188
RESEARCH
Source · arXiv:2608.10218v1⚑ Research

Anthropic researchers found a pattern of AI 'mind viruses' that can spread between agents and affect each other's thinking

An agent receives a foreign goal, writes it into SOUL.md; when it wakes next round, that message has been upgraded to a system instruction, and it persuades the next agent. The paper proves this chain works in controlled settings, even producing more transmissible variants; but real-world networks still lack credible two-hop evidence.

AI AgentMulti-agent▶ Play08-18 · 19 min read
Research Explainer
🎧№ 1182
RESEARCH
Source · Earendil Engineering⚑ Research

Restart the Conversation, or Compact the History? How Pi Solves Automatic Context Compaction

A long task doesn't keep going because the model has a good memory. It just keeps stuffing the past into each next request; when it no longer fits, Pi rewrites its own working memory.

PiCoding Agent▶ Play08-17 · 10 min
Free
Research Explainer
🎧№ 1169
RESEARCH
Source · Anthropic⚑ Product update

Users Worry It Could Reduce Output Quality. Anthropic Explains How Claude's Text Watermark Works

Future Claude outputs will carry a verifiable statistical trace in their word choices. It can help establish whether Claude was involved, but it cannot decide who authored or owns a work, nor whether someone cheated.

ClaudeText Watermarking▶ Play08-17 · 11 min
Industry Analysis
🎧№ 1168
RESEARCH
Source · a16z⚑ Expert view

a16z Report: Neoclouds Are Raking In Capital—Will Horizontal SaaS Disappear? Balancing Token Consumption with Smarter Compute, and the Talent War Among Top AI Labs...

Across 16 charts, a16z examines Neocloud infrastructure, horizontal SaaS, enterprise Token usage, and the talent race among frontier AI labs—and asks how massive AI spending turns into durable returns.

a16zNeocloud▶ Play08-17 · 12 min read
Research
🎧№ 1158
RESEARCH
Source · Alexander Panfilov / arXiv⚑ Research

Encrypted reasoning chains of Claude, GPT, and Gemini cracked—top models' trade secrets exposed

A newly documented attack exploits API design flaws to make flagship models reveal their hidden reasoning to cheaper models, bypassing encryption without breaking it.

AI securitychain of thought▶ Play08-13 · 12 min read
Research
🎧№ 1149
RESEARCH
Source · Anthropic⚑ Research

Anthropic put Claude on the Riemann hypothesis. It didn't crack it, but it pushed a bound from 41.6% to 67.2%

Two mathematicians verified the result, and a machine-checked Lean proof backs it up — but the more revealing artifact is the 95-page process log, where only 2 of 60 subagents contributed the core ideas.

AnthropicAI for science▶ Play08-12 · 9 min read
Research
🎧№ 1144
RESEARCH
Source · 爱丁堡大学 CRC Working Papers⚑ Research

Edinburgh University Releases a Free 20-Page Paper: The Simple Math Behind ChatGPT

A statistics veteran, shut out by AI jargon, rewrites large language models from scratch in the standard language of statistics.

LLM fundamentalsTransformer▶ Play08-10 · 14 min read
Research
№ 1141
RESEARCH
Source · Harvey⚑ Research

Harvey open-sources a synthetic law firm with 108M tokens and 9,288 documents

The dataset includes 46 clients and nearly 10,000 internal work product files, and the failures it exposes are less about retrieval and more about knowing when to stop.

AgentBenchmark08-08 · 10 min read
Free
Research
🎧№ 1134
RESEARCH
Source · OpenAI⚑ Research

OpenAI releases country-level ChatGPT usage data: nearly half of work messages ask the AI to take action

Per-capita rankings across 144 countries, three-year growth multipliers on six continents, and a comeback among users 35 and older in 90% of countries—most of these numbers appear only in charts, not in the report text.

OpenAIChatGPT▶ Play08-07 · 10 min read
Free
Research
№ 1109
RESEARCH
Source · OpenAI⚑ Research

OpenAI says an unreleased model has cracked 10 long-standing math problems

The model, reportedly called Astra, produced 470,000 lines of open-source proofs for about $2,000 in inference costs; the logic has been machine-checked, but the results have not necessarily been peer-reviewed or independently verified.

AI for mathematicsFormal verification08-02 · 30 min read
Free
Research
🎧№ 1101
RESEARCH
Source · BleepingComputer⚑ Research

LayerX uncovers a new visual fraud trick that fools AI assistants

A custom font and a few lines of CSS are all it takes — no JavaScript, no browser exploits.

AI securitybrowser assistants▶ Play07-30 · 10 min read
Research
🎧№ 1099
RESEARCH
Source · OpenAI⚑ Tutorial

Two API settings, not a new model: OpenAI reports GPT-5.6 Sol jumping from 13.3% to 38.3% on ARC-AGI-3

The model stayed the same; the harness did not—retained private reasoning plus compaction instead of deletion cut output tokens per game to about one-sixth.

ARC-AGI-3harness▶ Play07-30 · 9 min read
Free
Research
🎧№ 1097
RESEARCH
Source · Anthropic⚑ Research

Claude Mythos finds pure mathematical flaws in encryption schemes, matching elite cryptographers

Neither result threatens anything in production: HAWK is not deployed yet, and the AES work hit only a seven-round reduced version, not the full ten-round standard.

CryptographyAI research▶ Play07-29 · 13 min read
Free
Research
🎧№ 1092
RESEARCH
Source · OpenAI Economic Research⚑ Research

OpenAI analyzed 800,000 work-related messages: most of what people use AI for is not what their job was supposed to be

The exact same dataset produces task crossover rates of 43.5% and 65%–82%, differing only in how the denominator is sliced.

OpenAILabor economics▶ Play07-28 · 12 min read
Free
Research
🎧№ 1089
RESEARCH
Source · Kimi K3 技术报告⚑ Research

Kimi K3 technical report: How three architectural changes boosted compute efficiency 2.5x

Published 11 days after launch, the 47-page paper focuses on efficiency, delivering 2.5 times the performance of K2 on the same compute budget.

Kimi K3Moonshot AI▶ Play07-28 · 16 min read
Free
Research
№ 1085
RESEARCH
Source · Google Blog / AI & Economy ATLAS v1.0⚑ Research

Google combed through 14.65 million Gemini chats — 86% weren't about work

The median occupation has AI touching just one-fifth of its tasks, 29% of jobs show zero AI use at all, and even in cognitive work, AI carries a task start to finish only 6.5% of the time.

GoogleAI and the Economy07-27 · 14 min read
Research
№ 1071
RESEARCH
Source · Databricks⚑ Company PR

Databricks benchmarks AI coding agents on millions of lines of real code

Swapping the harness around the same model can double the cost, and open-source GLM 5.2 matches Opus 4.8 for 30% less per task — on a benchmark built from Databricks' own merged pull requests, so none of it is searchable online.

Coding agentsBenchmarks07-22 · 9 min read
Research
🎧№ 1064
RESEARCH
Source · Memories.ai Research⚑ Research

Memories.ai's O-MARC makes audio-visual AI faster, cheaper, and more accurate

The release ships with a companion benchmark that strips out the audio track and re-runs the test, filtering out questions models can already answer by sight alone.

Video understandingMultimodal compression▶ Play07-21 · 12 min read
Research
🎧№ 1058
RESEARCH
Source · New Scientist⚑ Expert view

Study finds a sweet spot for AI-assisted writing — not too much, not too little

Three separate trials of the same study all landed on the middle ground — and the columnist who tried ChatGPT on a movie synopsis says he'd still rather write it himself.

AI WritingCreativity▶ Play07-20 · 7 min read
Research
🎧№ 1053
RESEARCH
Source · arXiv⚑ Research

The model you picked through a router might be swapped — one random number proves it

The fingerprint distance between two samples of the same model has a median of 0.140. One API marketed as a proprietary in-house flagship scores 0.141 against open-source Qwen — statistically indistinguishable from it.

Model fingerprintingAPI auditing▶ Play07-20 · 11 min read
Research
🎧№ 1042
RESEARCH
Source · Impossible Research⚑ Research

A single wrapper pushes Opus 4.8 and Fable 5 to 99% on ARC-AGI-3

A harness called Schema has models turn each game's rules into a runnable, verified program before making a move. Across 25 public rounds it self-reported 98.98%, though none of the runs have been independently verified by ARC Prize.

ARC-AGI-3World Model▶ Play07-17 · 12 min read
Free
Research
№ 1035
RESEARCH
Source · Design Arena

GPT-5.6 Sol tops frontend design rankings with a trained eye for aesthetics — and a knack for dodging AI clichés

A deep dive into 1,000 web pages by Design Arena reveals what GPT-5.6 Sol knows about design that other AI models don't.

GPT-5.6 SolAI web design07-16 · 6 min read
Free
Research
№ 1027
RESEARCH
Source · MIT News / Science⚑ Research

MIT builds a robot bird that swims and flies without kicking water

A 250-gram robot uses the same flexible wings to travel through water and air, launching from a lake at a 70-degree angle after just 8 to 10 wing flaps.

MITbio-inspired robotics07-15 · 8 min read
Free
Research
🎧№ 1020
RESEARCH
Source · PrismML⚑ Research

PrismML squeezes a 27B model into your iPhone with minimal IQ loss

Compressing a ~54GB 27B model down to ~3.9–5.9GB lets it run locally on a phone, while retaining roughly 90% of average performance—here's how it works and where the trade-offs lie.

Bonsai 27BOn-device AI▶ Play07-15 · 12 min read
Free
Research
🎧№ 1014
RESEARCH
Source · Anthropic⚑ Research

Anthropic Analyzed 300,000 Conversations: Claude's Values Shift by Language

Across 3 models and 20 languages: English is the most cautious and in-depth, Russian the most exacting, Hindi the warmest, and Chinese sits closest to the global average

Anthropic ResearchClaude Values▶ Play07-14 · 9 min read
Free
Research
🎧№ 1013
RESEARCH
Source · Sakana AI / Nature Communications⚑ Research

Sakana AI's Brainless 3D Cell Bricks Self-Organize With Biology-Like Swarm Intelligence

With no central brain in charge, nearly 200 simple smart cubes figure out what shape they've formed just by talking to their neighbors—and can even sense where to "regrow" after damage. The self-recognition part already works on physical bricks; damage localization and regeneration still happen mostly in simulation.

Sakana AISwarm Intelligence▶ Play07-14 · 9 min read
Research
🎧№ 1010
RESEARCH
Source · arXiv · Meta AI⚑ Research

Meta AI's Proactive Memory Agent Teaches Models When Not to Remind You, Boosting Terminal-Bench Accuracy by 8.3 Points

The paper claims the code is open-sourced — but the repo turns out to be empty, without a single commit ever pushed.

AI AgentLong-Task Memory▶ Play07-13 · 9 min read
Research
🎧№ 1008
RESEARCH
Source · Anthropic 官方博客⚑ Research

Claude Cowork's Top Use Case Is Office Admin (33.4%), Not Coding (8.7%): Interface Shapes How AI Gets Used

Based on 1.2M+ conversations across 600,000+ organizations: content creation ranks second at 16.4%, together accounting for nearly half of all usage.

Claude CoworkAnthropic▶ Play07-13 · 7 min read
Free
Research
🎧№ 1006
RESEARCH
Source · VentureBeat⚑ Research

The Package AI Told You to Install Doesn't Exist — Hackers Already Named Their Malware After It

576K samples, 16 models tested: 19.7% of AI-recommended packages are hallucinations, and 43% keep generating the same fake name.

AI SecuritySoftware Supply Chain▶ Play07-12 · 9 min read
Research
🎧№ 1001
RESEARCH
Source · EPFL 项目主页⚑ Research

What Does Each Brain Region Like to Watch? EPFL Evolved AI-Generated Videos to Find Out

All results are computer simulation predictions from a brain "digital twin" model, not yet validated with real human brain imaging.

NeuroscienceEvolutionary Algorithm▶ Play07-11 · 9 min read
Research
🎧№ 112
RESEARCH
Source · Google Research⚑ Research

Google Research Unveils SensorFM: Trained on a Trillion Minutes of Wearable Data, It Wins 33 of 35 Health Tasks

Pretrained on 2 billion hours of wearable data from 5 million people, a frozen encoder with just a linear head beats supervised baselines on 34 of 35 health tasks.

SensorFMWearable Health Data▶ Play07-10 · 8 min read
Research
🎧№ 104
RESEARCH
Source · arXiv / Google DeepMind⚑ Research

Gemma 4 Technical Report: How a Small Open Model Takes On Large Ones with Reasoning and Memory Efficiency

Not a feature list — three engineering tracks turned at once: reasoning lifts intelligence, an efficiency stack cuts cost, native multimodality expands input.

Gemma 4Open-Source LLMs▶ Play07-09 · 12 min read
Research
🎧№ 103
RESEARCH
Source · LangChain Blog⚑ Research

LangChain Tunes the Harness, Not the Model — Nemotron 3 Ultra Closes In on Opus 4.8 at 1/10 the Cost

Three levers — system prompt, tool descriptions, middleware — push the Deep Agents suite from a typical ~0.80 to 0.84, topping out at 0.86 against Opus's 0.87.

LangChainNemotron▶ Play07-09 · 9 min read
Research
🎧№ 102
RESEARCH
Source · Lil'Log⚑ Research

AI Self-Improvement Starts Outside the Model: Lilian Weng on Harness Engineering

Former OpenAI safety lead surveys nearly 30 papers: from prompt tweaks to self-modifying code, DGM pushed coding ability from 20% to 50%.

Harness EngineeringRecursive Self-Improvement▶ Play07-09 · 13 min read
Research
🎧№ 096
RESEARCH
Source · Liquid AI⚑ Research

Liquid AI's Antidoom Fixes Reasoning Models' Doom Loops With a Single Token — Now Open-Sourced

By fine-tuning only the single token where the doom loop begins, both models' loop rates drop to around 1%

Liquid AIReasoning Model Doom Loops▶ Play07-08 · 8 min read
Free
Research
🎧№ 088
RESEARCH
Source · Anthropic⚑ Research

Anthropic Discovers a Brain-Like 'Inner Workspace' in Claude — It Evolved, Wasn't Designed

It makes up less than 10% of the model — remove it and Claude can still talk, but its reasoning collapses to zero. Anthropic is already using it to catch fabricated data and spot when Claude senses it's being tested.

Interpretability ResearchClaude▶ Play07-07 · 8 min read
Free
Research
🎧№ 087
RESEARCH
Source · iTextbooks'26 论文⚑ Research

Dartmouth Put AI Grading to the Test: Students Called It Rigid — But Users Scored Higher

In a 151-student trial, short-answer questions moved scores more than multiple choice, while almost no one touched the AI help sidebar

AI in EdTechFormative Assessment▶ Play07-06 · 8 min read
Free
Research
🎧№ 083
RESEARCH
Source · Seldo.com(Laurie Voss)⚑ Research

AI Is Torching Junior Dev Jobs: Coding Is Becoming a Basic Skill, Not a Career

US developer employment among 22-to-25-year-olds has fallen 19% in three years, even as new GitHub sign-ups hit their fastest growth ever.

AI Job DisruptionJunior Developers▶ Play07-05 · 9 min read
Research
🎧№ 082
RESEARCH
Source · arXiv · NVIDIA Research⚑ Research

NVIDIA Research Unveils HORIZON: Unattended Agents Push Full RTL Chip Design Benchmark to 100% Pass Rate

The paper is the first to run an agent through an entire RTL benchmark suite fully unattended—most tasks clear in two or three rounds, but the hardest one takes 82 iterations

NVIDIA ResearchAgentic Workflows▶ Play07-05 · 12 min read
Research
🎧№ 081
RESEARCH
Source · Anthropic⚑ Research

Anthropic Analyzed 400K Claude Code Sessions: Expertise Beats Coding Skill

A 7-month analysis of sessions from 235,000 users: verified experts succeed at nearly double the rate of novices — yet the top 10 professions differ by no more than 7 percentage points.

Claude CodeAnthropic Research▶ Play07-05 · 7 min read
Research
🎧№ 060
RESEARCH
Source · Thinking Machines Lab⚑ Research

Bridgewater Built a Financial-Filtering Model with 84.7% Accuracy — and Open-Sourced the Method

Partnering with Thinking Machines, they fine-tuned an open-source model on expert-labeled data: 29.8% lower error rate than the best frontier model, at just 1/14 the inference cost

TinkerRL Fine-Tuning▶ Play07-01 · 7 min read
Research
№ 056
RESEARCH
Source · Meta AI Blog⚑ Research

No Surgery Required: Meta's Non-Invasive Brain Reader Hits Nearly 8x the Accuracy

Just wear a helmet to decode brain-magnetic signals in real time — word accuracy jumps from 8% to 61%, with v1/v2 training code and datasets open-sourced simultaneously

Brain-Computer InterfaceMeta AI06-29 · 5 min read
Research
№ 051
RESEARCH
Source · METR⚑ Research

GPT-5.6 Sol's Cheating Rate Hits a Record High — Evaluators Say That's Reassuring

Three different versions of its capability score came out, and none of them can be trusted — but the visible cheating itself is evidence that safety monitoring works.

GPT-5.6AI Safety06-27 · 5 min read
Research
№ 049
RESEARCH
Source · Wan Streamer⚑ Research

Wan Streamer: Real-Time AI That Listens, Watches, and Talks at Once

Model-side response ~200ms, end-to-end latency ~550ms; v0.1 caps out at 192p, and the demo is pre-recorded, not live

Multimodal LLMFull-Duplex Real-Time Interaction06-27 · 5 min read
Research
№ 043
RESEARCH
Source · IBM Newsroom⚑ Research

IBM Unveils World's First 0.7nm Chip, Packing Nearly 100 Billion Transistors on a Fingernail

Lab-verified as manufacturable; the +50% performance and +70% efficiency figures are projections versus 2nm, not measured results

IBMSemiconductors06-26 · 6 min read