Most companies run plenty of AI pilots, but they never change an end-to-end business process, and they never turn a one-off win into a capability they can reuse next time. The dividing line is whether the company ends up with its own data, its own ways of working, and its own record of what it has learned.
A single answer is no longer just a broad sweep of the web. ChatGPT may now discover candidate sources first, then lock onto specific sites to dig deeper. As the retrieval order shifts, GEO expands from page-level ranking into a three-layer contest: domain inclusion, page matching, and citation conversion.
Moderna and Merck's personalized mRNA cancer vaccine hit positive results in a Phase 3 trial for melanoma. Algorithms help select each patient's tumor neoantigens, then a custom vaccine is manufactured for that individual. This is the first time a truly personalized approach has cleared a large-scale clinical hurdle, but widespread use still faces three big challenges: data, cross-cancer efficacy, and cost.
From Git's DAG and packfiles to GitHub Spokes, and then to Continuity with S3 WAL as the source of truth: a full walkthrough of why large-scale Git hosting is hard and how Cursor solves it.
An agent receives a foreign goal, writes it into SOUL.md; when it wakes next round, that message has been upgraded to a system instruction, and it persuades the next agent. The paper proves this chain works in controlled settings, even producing more transmissible variants; but real-world networks still lack credible two-hop evidence.
A long task doesn't keep going because the model has a good memory. It just keeps stuffing the past into each next request; when it no longer fits, Pi rewrites its own working memory.
Future Claude outputs will carry a verifiable statistical trace in their word choices. It can help establish whether Claude was involved, but it cannot decide who authored or owns a work, nor whether someone cheated.
Across 16 charts, a16z examines Neocloud infrastructure, horizontal SaaS, enterprise Token usage, and the talent race among frontier AI labs—and asks how massive AI spending turns into durable returns.
A newly documented attack exploits API design flaws to make flagship models reveal their hidden reasoning to cheaper models, bypassing encryption without breaking it.
Two mathematicians verified the result, and a machine-checked Lean proof backs it up — but the more revealing artifact is the 95-page process log, where only 2 of 60 subagents contributed the core ideas.
A statistics veteran, shut out by AI jargon, rewrites large language models from scratch in the standard language of statistics.
The dataset includes 46 clients and nearly 10,000 internal work product files, and the failures it exposes are less about retrieval and more about knowing when to stop.
Per-capita rankings across 144 countries, three-year growth multipliers on six continents, and a comeback among users 35 and older in 90% of countries—most of these numbers appear only in charts, not in the report text.
The model, reportedly called Astra, produced 470,000 lines of open-source proofs for about $2,000 in inference costs; the logic has been machine-checked, but the results have not necessarily been peer-reviewed or independently verified.
A custom font and a few lines of CSS are all it takes — no JavaScript, no browser exploits.
The model stayed the same; the harness did not—retained private reasoning plus compaction instead of deletion cut output tokens per game to about one-sixth.
Neither result threatens anything in production: HAWK is not deployed yet, and the AES work hit only a seven-round reduced version, not the full ten-round standard.
The exact same dataset produces task crossover rates of 43.5% and 65%–82%, differing only in how the denominator is sliced.
Published 11 days after launch, the 47-page paper focuses on efficiency, delivering 2.5 times the performance of K2 on the same compute budget.
The median occupation has AI touching just one-fifth of its tasks, 29% of jobs show zero AI use at all, and even in cognitive work, AI carries a task start to finish only 6.5% of the time.
Swapping the harness around the same model can double the cost, and open-source GLM 5.2 matches Opus 4.8 for 30% less per task — on a benchmark built from Databricks' own merged pull requests, so none of it is searchable online.
The release ships with a companion benchmark that strips out the audio track and re-runs the test, filtering out questions models can already answer by sight alone.
Three separate trials of the same study all landed on the middle ground — and the columnist who tried ChatGPT on a movie synopsis says he'd still rather write it himself.
The fingerprint distance between two samples of the same model has a median of 0.140. One API marketed as a proprietary in-house flagship scores 0.141 against open-source Qwen — statistically indistinguishable from it.
A harness called Schema has models turn each game's rules into a runnable, verified program before making a move. Across 25 public rounds it self-reported 98.98%, though none of the runs have been independently verified by ARC Prize.
A deep dive into 1,000 web pages by Design Arena reveals what GPT-5.6 Sol knows about design that other AI models don't.
A 250-gram robot uses the same flexible wings to travel through water and air, launching from a lake at a 70-degree angle after just 8 to 10 wing flaps.
Compressing a ~54GB 27B model down to ~3.9–5.9GB lets it run locally on a phone, while retaining roughly 90% of average performance—here's how it works and where the trade-offs lie.
Across 3 models and 20 languages: English is the most cautious and in-depth, Russian the most exacting, Hindi the warmest, and Chinese sits closest to the global average
With no central brain in charge, nearly 200 simple smart cubes figure out what shape they've formed just by talking to their neighbors—and can even sense where to "regrow" after damage. The self-recognition part already works on physical bricks; damage localization and regeneration still happen mostly in simulation.
The paper claims the code is open-sourced — but the repo turns out to be empty, without a single commit ever pushed.
Based on 1.2M+ conversations across 600,000+ organizations: content creation ranks second at 16.4%, together accounting for nearly half of all usage.
576K samples, 16 models tested: 19.7% of AI-recommended packages are hallucinations, and 43% keep generating the same fake name.
All results are computer simulation predictions from a brain "digital twin" model, not yet validated with real human brain imaging.
Pretrained on 2 billion hours of wearable data from 5 million people, a frozen encoder with just a linear head beats supervised baselines on 34 of 35 health tasks.
Not a feature list — three engineering tracks turned at once: reasoning lifts intelligence, an efficiency stack cuts cost, native multimodality expands input.
Three levers — system prompt, tool descriptions, middleware — push the Deep Agents suite from a typical ~0.80 to 0.84, topping out at 0.86 against Opus's 0.87.
Former OpenAI safety lead surveys nearly 30 papers: from prompt tweaks to self-modifying code, DGM pushed coding ability from 20% to 50%.
By fine-tuning only the single token where the doom loop begins, both models' loop rates drop to around 1%
It makes up less than 10% of the model — remove it and Claude can still talk, but its reasoning collapses to zero. Anthropic is already using it to catch fabricated data and spot when Claude senses it's being tested.
In a 151-student trial, short-answer questions moved scores more than multiple choice, while almost no one touched the AI help sidebar
US developer employment among 22-to-25-year-olds has fallen 19% in three years, even as new GitHub sign-ups hit their fastest growth ever.
The paper is the first to run an agent through an entire RTL benchmark suite fully unattended—most tasks clear in two or three rounds, but the hardest one takes 82 iterations
A 7-month analysis of sessions from 235,000 users: verified experts succeed at nearly double the rate of novices — yet the top 10 professions differ by no more than 7 percentage points.
Partnering with Thinking Machines, they fine-tuned an open-source model on expert-labeled data: 29.8% lower error rate than the best frontier model, at just 1/14 the inference cost
Just wear a helmet to decode brain-magnetic signals in real time — word accuracy jumps from 8% to 61%, with v1/v2 training code and datasets open-sourced simultaneously
Three different versions of its capability score came out, and none of them can be trusted — but the visible cheating itself is evidence that safety monitoring works.
Model-side response ~200ms, end-to-end latency ~550ms; v0.1 caps out at 192p, and the demo is pre-recorded, not live
Lab-verified as manufacturable; the +50% performance and +70% efficiency figures are projections versus 2nm, not measured results