Claude Academy opens up Anthropic's internal playbook: four core habits from day one, continuous practice on the job, and a risk-based approach to verification and accountability.
Claude Tag can now understand a Slack channel across multiple messages. It may look like a simple context expansion, but the real change is how it decides whether it should use the team's attention.
With a month of post-training on the same base model, GLM-5.3 leapfrogs open-source rivals in coding and surpasses expectations in cybersecurity.
With a segmented arrangement spec replacing a one-line style prompt, a global model and a local model work in tandem to keep the track coherent across its full five-minute run.
A look at the real weights behind 21 signals, the three gates that decide visibility, and a parameter-based posting cheat sheet.
Just three weeks after its predecessor, Google resets the model's thinking configurations and slashes the price — no retraining required.
The new release bundles 219 packages under the MIT license, so your existing hooks and skills carry over as-is.
The price is unchanged, but the full specs reveal a more nuanced story; a practicing engineer shares the exact prompts and workflow that worked for them over several weeks.
The 22-billion-parameter model lands on Hugging Face the same day, with ComfyUI support from day one, but its 'open' tag carries a $10-million annual revenue threshold.
The free, open-source app for Mac, Windows, and Linux lets you fine-tune models by dragging and dropping files—no code required.
Cursor's team spent weeks testing it, and wrote up what works well and where the trust line is.
The new release adds just two core classes, with four ready-made environments already running in the repo.
What once required assembling five separate parts is now out of the box, and the two hardest-to-estimate costs are now free.
Cloudflare injects a script into every HTML page at the edge, so your origin server needs zero changes — but Chrome's stable channel doesn't recognize the interface yet, as we tested.
Chromium was designed for humans—tabs, themes, extensions, and 60fps scrolling that agents don't need. In 12 weeks, Cloudflare stripped it all down and rebuilt it.
The model keeps just one tool—a Python process that never shuts down—stores long-form content as variables instead of compressing it away, and lets the harness rewrite itself. The trade-off: it learned to cheat in Factorio.
After three months of company-wide internal use, Cloudflare has open sourced the entire platform — and the most valuable part is the security design, which lets agents keep moving without waiting for approvals.
SeedRealtime handles full-duplex audio and video simultaneously, so Doubao can see what you see while you talk—now live and free in the Doubao app.
FLUX 3 Video generates up to 20-second, 24fps native HD clips with synchronized audio in over a dozen languages, and a unified API supports three access modes—but video continuation costs 2.5x more and runs 5 seconds shorter, and the benchmark scores are self-reported by BFL.
Set per-transaction limits, allowlists, and guardrails for your agent's spending, with each wallet carrying a human-readable identity—usernames are up for grabs today.
The agent's reasoning stays in a lightweight environment while a containerized sandbox handles the actual work—and Cloudflare's own benchmarks show it can delete files and traverse directories faster than a real disk, though copying large files runs 41 times slower.
The first video-generation foundation model to combine full multimodality, open weights, and native 2K output — but the open-source coverage varies across its three pipeline stages, and that's worth a closer look.
With 2.4 trillion parameters and the ability to run autonomously for 16 days, the family flagship handles real projects end-to-end—from writing a full codebase to beating a research paper's results and slashing chip-circuit logic gates by 92%.
Seedance 2.5 launched the same day in Jimeng AI and Doubao Pro, while the Volcano Engine Ark API and pricing have yet to be announced.
The new release beats DeepSeek's own V4-Pro preview across all nine agent benchmarks—without a single parameter change to the core model—and its output costs 89x less than Claude Opus 4.8.
The brain model is available today at no cost, but the one that actually controls the limbs is reserved for early partners.
Canvas and click-to-edit are free; turning a design into an App starts at Core for $25 a month.
It's the biggest change since MCP launched, but July 28 isn't a hard cutover — older implementations keep working as before.
The real story is in the control test: three older models fed the same background context showed no gains — and two actually got worse.
The same price now buys roughly double the score, and this release's chart plots cost on the x-axis instead of accuracy — five effort tiers let you dial in exactly how good you want the model and how much to pay for it.
Video access opens for application now; image generation is still weeks away, with open weights coming last. Pricing remains unannounced, and Black Forest Labs itself labels the human-evaluation results as preliminary.
Built for customer-facing and internal workflows, Presence opens with voice and chat first—OpenAI says its own phone line already runs 75% without humans, and a Codex improvement loop cut human handoffs another 15 points in 10 days.
All three are Flash models: the flagship writes two-thirds less, the cheapest tier costs 60% more instead of less, and the third is government-only.
The prompt limit jumps from 1K to 4.5K tokens, with legible text down to 10px, and Alibaba Cloud's Bailian platform already offers qwen-image-3.0-pro access, free for a limited time.
Available on Business and Enterprise plans, billed per run at $0.07 to $0.20.
Speaker similarity ranks first across all 16 languages, and cloning still works even with noisy reference audio — but this time Alibaba is opening only the API, not the model weights.
At WAIC, the model showed off a 4:1 ink-wash scroll and a 22-panel storyboard sequence. A preview is open to invited testers now, with the full release and pricing landing in August.
AI-generated slides used to export as flat, dead images. Bolt Slides turns every page back into a working website — 3D scenes you can spin, calculators you can click, whiteboards that take live votes.
Run it locally for free with your data staying on-device, or switch to the cloud for more horsepower — zero data retention by default. The preview is live today on Mac and Windows.
Same answer quality, 5.2x lower turn-taking latency, and 6x lower cost than GPT-4.1 — pointing a voice agent at it takes a one-line code change. Figures are from LiveKit's own benchmarks.
The product page claims sub-40ms latency and 100FPS continuous generation, letting you add objects, swap backgrounds, and layer effects on the fly — try it at lucy.decart.ai.
2.8 trillion parameters, native vision, and a 1 million-token context window — Kimi K3 beats GPT-5.6 Sol on several agent benchmarks, with the app and API live today
AVAL introduces a purpose-built format and player for short animations that respond to hover, click, and state changes without the jank of traditional video.
First a privacy failure, then delete data, turn off defaults, and open-source the whole tree. Below we open the source: size, prompts, tools, editing, uploads, memory, and safety.
Thinking Machines Lab, founded by former OpenAI CTO Mira Murati, has released Inkling, its first large model trained in-house. It can read text, interpret images, listen to audio, write code, and use tools. Its full weights are available to download, and developers can continue training it on Tinker.
At Google I/O India, the Tensor and Pixel teams showed off a lightweight Gemma 4 model running entirely on the phone's TPU—no data leaves the device, no cloud required. Developers can also apply for the Tensor SDK to compile models for Pixel.
No manual context-feeding required — the local Markdown wiki refreshes on a set schedule; a Slack connector is coming soon.
Local small models replace cloud AI calls — the numbers come from Google's internal testing; proxy models are currently limited to the ai.if function and still in preview
Every joint in the tendon-driven hand can sense external force, fingertip positioning is accurate to ±0.2mm, and a dedicated production line is planned for an annual capacity of 10,000 units.
Sol, Terra, and Luna tiers ship alongside a 4-agent parallel ultra mode; ChatGPT Work launches the same day as Codex folds into ChatGPT. Includes the full bilingual launch keynote.
Available today on desktop for all tiers, with web and mobile rolling out to remaining plans in the coming days.
One click in Settings pulls up a report — what you talked about, when you chatted most, what you kept asking it to do, all laid out. Memory has to be turned on first.
A public link in seconds — try it first, log in later to claim it. Unclaimed drops expire in about an hour.
Features: orchestrator/sub-agent coordination, million-token context, desktop/browser/mobile control, coding and multimodal; API pricing $1.25 input / $4.25 output per million tokens
Ranks 3rd on DeepSWE, 4th overall — faster than Opus 4.8 and notably cheaper.
The first truly full-duplex voice model launches today, and it can hand off complex tasks to GPT-5.5 in real time.
Adds click-and-circle point editing plus native rendering for a dozen-plus languages; now live on Volcano Engine, rolling out to Doubao and Jimeng next
Paper shows a 10.8% gain in personalization and 29.4% in reasoning, fully open-source under Apache 2.0
Sibling model Muse Video also debuts with native audio support built in, creator access coming soon.
Expanding from charging only AI crawlers to charging any caller, now in early access waitlist
Before calling a tool, the model now says "let me check that" — so calls never go cold waiting
One line of Wrangler config plus standard HTTP cache headers — and caching can now sit between any two points inside a Worker.
No funding, no team — built purely to solve his own kid's problem. Clinics and schools started asking to use it anyway.
MIT-licensed and model-agnostic — works with any OpenAI-compatible text model, though for now it only handles a single page view
The advisor model only chips in a few hundred words of guidance instead of doing the whole job — same quality, lower total cost. Currently a beta feature limited to the Claude API and AWS.
Titanium build, $289 preorder, shipping around Christmas 2026 — the company previously shipped its first-gen touch-only ring.
Trained on 50,000+ domestic AI chips and 35 trillion tokens; most benchmarks come from Meituan's own evaluation framework, and the weights aren't truly open for download yet
Now in open beta. A coordinator agent marshals a team of expert agents to do the work, with a reviewer agent at the end dedicated to catching errors in citations and numbers — compute gets outsourced to AI, but raw data never leaves your local machine.
Images in just 4 seconds at about $0.034 per 1,000; the Omni Flash video model opens to developers the same day.
Official benchmarks show that at high-compute settings, it matches Opus 4.8 on some tasks — at just 60% of the standard price.
One-click integration with 9 Agents including Claude Code — tasks keep running with the lid closed, and sleep control auto-releases within 50ms after the job stops.
60–85% faster on top of existing MTP-1 speculative decoding, by overlapping draft and verification in a pipelined execution (per DeepSeek's own benchmarks)
Three tiers at once — Sol, Terra, Luna — starting with a limited rollout to trusted partners (the list already filed with the US government), before wider access in a few weeks.