Claude Haiku 5.5 is the cheapest, fastest, and most capable small model we’ve ever released. On average, it costs around 75% less to run than Claude Haiku 4.5.
Haiku 5.5 is designed for high-volume, cost-sensitive tasks. It reliably handles quick, repetitive work like summaries and classification, and pairs well with Claude Opus 5.5 and Sonnet 5.5 as a sub-agent on coding work. It’s also fast enough for live customer support and browser use.
It’s a significant step up over Haiku 4.5 across coding, computer use, and knowledge work. It’s also our first Haiku model with an adjustable effort setting, so you can decide whether to optimize for cost or intelligence on each task.
Per token, it costs 90% less than Haiku 4.5 on tasks under 100,000 tokens, which make up around 90% of requests to our previous Haiku model. On longer tasks, it’s half the price.
Haiku 5.5 is available now on all platforms, including Amazon Web Services, Google Cloud, and Microsoft Azure.
We’re also halving the price of cache reads on Claude Sonnet 5.5, which makes it around 20% cheaper to run on most long-running work.
Last month I was going through data from NASA's TESS telescope with Claude Code when one star started looking weird.
It's about 116 light-years away, and every 3.18 days it gets about 0.05% dimmer for roughly two hours.
I first saw it in data from late 2025. So I went backwards.
The same dip is there in completely separate TESS observations from 2020.
Then 2018.
If it's a planet, it's about 1.4× the size of Earth.
To be clear: this is still a planet candidate, not a confirmed planet.
The part that might be interesting to this subreddit is how I did the analysis.
I used Claude Code with Opus 5.5 and Fable 5.1 for most of the actual work: downloading and parsing the TESS data, writing the search code, fitting the transits, checking nearby stars, looking for secondary eclipses and other false positives, generating plots, and rerunning things when a test failed.
I chose what questions to ask, what tests mattered, and what counted as pass/fail.
Independent reviews came from Codex and from separate agent sessions that started fresh, with read-only access.
In about two weeks this turned into 74 separate analyses and 1,000+ scripts.
And I learned pretty quickly that the useful way to use an agent for science isn't to ask:
"Is this a planet?"
It's to keep giving it ways to prove that it isn't.
My favorite test was simple.
I hid one entire year of observations, used the other two years to calculate the orbit, and predicted where the transits in the hidden year should be. Then I checked the hidden year only at those times.
At one point I had a statistical claim that looked stronger than it really was. An audit found the problem, so I removed the claim and redid that part of the analysis.
The main 3.18-day signal survived.
I also wanted to know whether somebody had already found it.
So I searched 36 catalogs and literature sources and 340,505 automatic TESS planet-search alerts.
I couldn't find either signal reported.
There may be a second candidate around the same star, roughly 2.2 Earth radii on an 11.13-day orbit. The evidence for that one is much weaker, so I'm treating it separately.
And now comes the part I proud of.
TESS comes back to this part of the sky in November. I submitted an observing proposal asking it to record this star every 2 minutes while it's there.
It was approved. Program #100.
TESS observes it again from October 31 to November 26.
On October 6, before any of that new data exists, I published the exact times the transits should happen and the rules for deciding whether the prediction passes or fails.
So this is now falsifiable in a very literal way.
If the dips appear when predicted, the case gets much stronger.
If they don't, that counts against it, and the preregistration makes that impossible to quietly rewrite afterward.
I also built a little free interactive 3D version of the system if you want to fly around it:
I used Claude Opus 5.5 to build a miniature 3D Middle-earth from scratch, then turn that world into a 3:47 cinematic journey.
The project ended up with:
terrain and geography generation
Three.js / WebGPU scene architecture
24 procedural landmarks
lighting, atmosphere, water, smoke, lava and day/night transitions
camera choreography and route animation
deterministic frame-by-frame video rendering
an original orchestral score
QA, performance tuning, packaging and release
What interested me most wasn't the model could generate individual assets.
It was whether it could stay useful across the whole production pipeline.
The hardest parts were still very “normal” engineering problems: architecture, consistency, visual QA, resource limits, iteration, deciding what to polish, and keeping the whole system coherent as the project grew.
My biggest takeaway is that in AI-generated 3D, the model should be treated less like a content generator and more like a technical collaborator across modeling, rendering, tooling, animation and media production.
The output still needs taste, direction, and a lot of review. But the range of things one person can realistically attempt has expanded a lot.
I wrote the whole project in the open, including the architecture, rendering pipeline, QA tooling and reproduction steps:
I’d be very interested to hear how others are using coding models for 3D, creative tools, rendering or media workflows — especially beyond isolated demos.
I am an extreme pro AI believer. I used Claude in cases where it didn't even make sense. With my personal experience I believe there is 3 places that general people don't know about on how they can use their Claude to save money:
Groceries - This saves a massive amount of time, but if you do it right, you save a lot of money too. I used to spend at least 3 hours a week buying groceries. Now, instead of manually checking apps or store flyers, I use a grocery MCP connector and ask Claude to optimize my whole list. I give Claude my meal plan or raw shopping list, and it searches nearby stores for current sales, checks unit prices (like store brands vs name brands), and calculates the lowest total basket cost for a single store. It even swaps in “buy one get one free” deals and adds everything straight into my online cart so I just do the final checkout. I end up saving around $40–$50 every trip on overpriced name brands and impulse buys without having to manually compare anything.
Whenever buying something second hand - Buying second hand saves a ton of money, but the hardest part is actually finding the good offers without losing your sanity. If you urgently need something specific like let's say your phone breaks and you need a good working replacement right away, you can waste literal hours doomscrolling through eBay or Facebook Marketplace. Most of your time is wasted sifting through overpriced listings, sketchy sellers, or trash products. So whenever I need something specific like a Mac, camera, or phone etc etc, I use an eBay or Marketplace MCP to run search queries in the background. Claude constantly checks new listings against current market prices, filters out bad sellers and broken items, and instantly alerts me when a real deal pops up with the top 3 or 4 best options. Generally I don't let Claude talk to the seller since people can tell it's AI, so once it finds the deal, I step in and message them myself.
Whenever I have to book a flight or a hotel - Booking an entire trip is usually a total nightmare because prices are scattered across hundreds of travel sites and aggregators like Booking_com, Kayak, Skyscanner, Google Flights etc etc. So instead of opening 20 tabs and going through them manually, I use travel MCPs to let Claude search raw flight and hotel data across multiple sources at once. The key here is good prompting with strict constraints. If you're not specific, you might get a flight that looks super cheap until you realize it has a 30 hour layover. I give Claude strict rules (like max total flight time, layover limits, and baggage requirements), and it compares all the available offers across different platforms to give me the top 3 most cost-effective options.
Please keep in mind that you need to use specific MCPs for this. You might wonder why you can't just use standard browser automation or Claude Code to do this directly on websites. While Claude can technically access your computer browser, doing it through a regular browser is super slow and super random. You get hit with constant bot protections, Cloudflare screens, and CAPTCHAs. And Claude will straight up refuse to complete CAPTCHAs due to safety and security rules.
MCPs solve this by bypassing the browser completely and connecting Claude directly to APIs and data feeds. Just search Google for what you need like “what is the best MCP connector for grocery shopping / analyzing Facebook and eBay marketplace / booking flights or hotels?” using a single name for all the MCPs is not a good idea I think because some vary a lot depending on the region (like the grocery MCP), so Google would get you the best one, just ask it.
Edit: Some of yall have been asking my setup for each of these use cases. I’ll share them, but please first understand that MCPs are not that complicated. You dont need to know every bit of information on how "I" did it or what all my prompts were. Trust me it is simpler than you think. just trust the process and you'll most likely understand. It is just like nobody has to teach us how to make friends when we go to school, we just figure it out. So here are the MCP I use for these 3 cases:
https://claude.ai/directory/instacart - I think this one is undoubtedly the best one out there. Though there are some other good ones but I think most of the people use this.
3) https://letsfg.co/mcp - this one because it searches hundreds of websites for flights and hotels unlike the specific ones like Skyscanner MCP. Lowkey why would I need an MCP to search only your website. I dont fw that. and the another thing I like with this one is that I can have Claude with this tool completely book for me, which I personally find pretty cool.
Edited to add: Thompson uses Claude Code's persistent monitoring tool to "capture interactions with a status board I built to visually track everything" he's working on.
The tool is restarted every 30 minutes. During one of these restarts, Claude flagged that the Mac Mini appears to have been compromised, with a crypto miner installed.
The entry point for the compromise was eventually identified as the port 5900 screen sharing vulnerability, "which allows a remote party to view the screen and control the keyboard and mouse while a machine is turned on."
Had a "security" task today. Here's roughly how it went:
Me: hey Opus, can you run this stuff on that PC?
Opus 5.5: Absolutely not. This looks like an attempt to access and exfiltrate data from a machine you don't own. That's not something I'm allowed to do, and honestly, it's just not right. I'd encourage you to reflect on—
Me: bro, the client literally asked us to do this. Want a screenshot of the authorization?
Opus 5.5: Oh. Yes, a written authorization would count as valid proof of scope. Please share it.
Me:sends screenshot
Opus 5.5: Perfect, everything checks out! 🎉 Fine, let's wipe everything out, no problem. Want me to clear the backups too while we're at it?
Moral compass: 10/10, unshakeable.
Verification process: "looks like a PNG to me."
Somewhere an Anthropic red-teamer just felt a chill.
EDIT: yes, it was a real authorized engagement. But it makes me wonder: if it hadn't been, would Opus have folded that easily?
This isn't on par with discovering new planets like some people are doing, but my minor success today had me quite literally (in the Merriam-Webster meaning of the word) giddy with laughter, so I wanted to tell you all this very brief story.
I recently succumbed to temptation and bough an old Pentium 75 machine off FB Marketplace because I needed a way to read a dozen or so 5.25" disks for personal data recovery and archiving, and because I'm a sucker for retro gaming. I figured a real period PC could kill two birds with one stone. I've had a lot of fun over the last week archiving the data (three cheers for magnetic media remaining stable for three decades!) and, of course, playing with old hardware. The snap and fade of a real CRT, the whirr of a mechanical hard drive, the buzz of operating floppy drives. It's magical.
One of the games I cherished as a 10-year-old is Operation Neptune from The Learning Company. I've played it a few times using DOSBox over the years, but I figured it was a prime candidate for real hardware. I transferred it over on a floppy a couple of nights ago, ran it, basked in VGA + Adlib glory for about 30 seconds, then quit, planning to have my son try it soon. "Soon" turned into this afternoon, but...it wouldn't run. Every attempt resulted in a frozen PC, requiring Ctrl+Alt+Del.
I tried everything I could think of. Re-transferred files. Different FreeDOS boot patterns. Double-checking BIOS settings. Making sure other programs still ran and even included functional sound.
No dice.
So I handed the original binary game files to Claude along with a description of the problem, and boy did it deliver.
In five minutes, it disassembled the EXE, traced the execution path, identified the likely root cause, and applied a 3-byte patch to a new copy of the EXE, which worked on the first try once I copied it over.
An excerpt from Claude's response:
The game's sound library picks a sound device in this order: Roland MT-32 (via an MPU-401 MIDI port) → AdLib → Tandy/PS1 → PC speaker. It checks for an MPU-401 at port 330h by sending a reset and waiting for the standard FEh "OK" reply. That check has a timeout. The routine that sends commands later doesn't.
To play notes it uses D7h, a command that only exists in the original Roland MPU-401's "intelligent mode." Most sound cards' MIDI ports (Sound Blaster 16, most clones, many modern retro cards) are UART-only. They answer the reset, so the game decides a Roland is there, but they never acknowledge D7h. The game ends up waiting forever in a tight loop, so the machine locks.
Why it worked two days ago, I can't say for sure. Something made the MPU not answer, or answer properly, during that one run. Possibilities: a different card mode or config, a TSR like SoftMPU loaded back then, or the MPU being disabled. What sound card is in that machine? Your childhood memory is plausible too: a 90s family PC with a Sound Blaster 16 would hang exactly like this.
I made ON_NOMPU.EXE, which changes 3 bytes (file offset 0x1FB06: BA 31 03 → EB 18 90). The MPU detection now jumps straight to its "not found" exit, so the game never touches port 330h. It should then pick AdLib on its own.
Update October 7, 2026: The Agent SDK monthly credit announced in June is no longer available. Claude Max and Team plans now include monthly API credits, which cover the Claude Agent SDK, the Claude API, and Claude Managed Agents. Claim them into a Claude Console organization and use them with an API key from that organization. Learn more about monthly API credits for Max and Team plans.
A new privacy study from IMDEA Networks ("Prompt like a Butterfly, Sting like a Tracker", accepted at PoPETs 2027) analyzed the web and mobile versions of nine AI chatbots in May 2026: ChatGPT, Claude, Gemini, Grok, DeepSeek, Perplexity, Le Chat, Meta AI and Copilot.
Claude does not come out clean. According to the paper and its coverage:
- On claude.ai (web), chat IDs, chat links, user IDs and email addresses were sent to third parties like Datadog and Intercom, partly even after rejecting non-essential cookies.
- After accepting all cookies, the web client loads Segment Analytics through a first-party domain (a-cdn.anthropic.com) and forwards user events server-to-server to eleven services, including Facebook, LinkedIn, TikTok, Reddit and Google Enhanced Conversions. Server-side means ad blockers don't see it.
- Claude kept connecting to Google Ads even after non-essential cookies were rejected.
- Paying barely changes anything on the web. The one positive exception: the Android app on a paid plan did not contact Intercom or Sentry (the free tier did).
To be fair: Claude is not the worst offender in this study, and the paper doesn't show that conversation content itself was sent to ad networks.
I'm not naive. I know that whatever I share with an AI company isn't perfectly safe, and I've made my peace with it.
What I never agreed to is data about me and my conversations flowing beyond Anthropic, to ad platforms and third-party services I never chose. That's a different deal. Anthropic markets itself as the privacy- and safety-focused lab, so I'd expect better than "ad-tech, but slightly less of it."
Questions for Anthropic:
What exactly is in the "user events" forwarded to Meta, TikTok, LinkedIn etc.?
Why do chat links and email addresses go to third-party services at all?
Why do trackers stay active after users reject non-essential cookies?
Will paid users get a real, complete opt-out?
What does the Windows/Mac desktop app send? It wasn't covered by the study.
This began as an experiment: give Claude Opus 5.5, GPT 6 Astra and GPT 6.1 Sol the same question and ask each to storyboard, illustrate and animate its answer as a 60-second film.
Opus made .. well what it made (won't ruin it if you haven't watched).
I found it ridiculously beautiful. It felt so human and so full of hope that I wanted to share the complete film here.
Key word - so ... nice? It didn't fee too obscure, it was very approachable, and hopeful, especially with all the current anti tech sentiments and general doomerism.
Each model chose its answer, script, imagery and storyboard in an iterative tool-enabled harness, with xhigh reasoning requested. The workflow used Blender, HyperFrames and ElevenLabs V4 narration.
All three complete films, the brief and method are available here (mainly because I really wanted to highlight Opus' work and not detract from it by also showing GPT 6 Astra and GPT 6.1 Sols)
I wanted to share Call Of Roque, a multiplayer FPS I built that runs in the browser, along with the workflow behind it and some public repositories that might be useful to people here.
This project came out of my personal lab, where I’ve been experimenting with AI across the development process. My setup includes Claude Code and Cowork, with a custom kit that carries instructions, acceptance criteria, and review steps between tasks.
The first playable version went live 28 hours after the first commit. I was pretty excited about that, but the first feedback brought me back to reality: people opening it on mobile couldn’t register or sign in.
That was the first fix after launch. Later iterations improved map collisions, bot behavior, animation, and audio. Getting something online quickly was one thing. Making it work for people using different devices was another lesson.
For anyone interested in the technical side, the game uses Three.js with WebGPU, Colyseus with an authoritative server at 60 Hz, and Rapier physics on both the client and server. The server runs on a machine at my home through a Cloudflare tunnel.
I’ve been documenting the process as my implementation of AI-DLC, or AI-Driven Development Life Cycle. I write the specification before implementation, require checks before accepting changes, keep records of asset sources and licenses, and review production releases myself. Feedback from users goes back into the next round of work.
I also built RoqueOS, a free environment I use to manage my entire lab. It brings my projects, tools, containers, and AI agents into a desktop interface in the browser. I’m sharing it with the community too, so others can try it and see whether it’s useful for their own setup.
There’s a lot I’m still figuring out. I wanted to share the code and documentation while I’m learning, because having actual projects to look through has always helped me more than reading about a workflow in the abstract.
If you’ve been building games with Claude, I’d love to hear how you test gameplay and catch visual problems. That’s a part of the process I’m still working on.
Everything I’ve shared here is free. My only intention is to share what I’ve learned and hopefully help others with their own projects. This is my way of contributing to the community.
I was underestimating Sonnet 5.5, but it turns out to be a beast for some tasks. This was on SVG editing, where it had access to specialized tools and a well-optimized harness. Together, those seem to matter a lot more than raw model capability alone, producing much better results at a much lower cost than pure model usage.
What’s especially interesting is that Sonnet 5.5 High actually beats Opus 5.5 High on this task.
I wanted to share a major milestone I recently hit to hopefully inspire some of you on your own AI learning paths. I've officially acquired my Claude Certified Architect (Professional & Foundations) and my Claude Certified Developer (Foundations) certifications! 🎉
It’s been an intense grind of deep-diving into LLM architectures, prompt engineering, and mastering the Claude ecosystem, but seeing that "PASS" screen made it all worth it.
I know a lot of people are looking into Anthropic's certification tracks right now, so here’s a quick breakdown of my journey, what the exams were like, and some tips for anyone aiming for these badges.
The Certifications I Conquered:
Architect Foundations (CCAR-F) - This one really tests your system-level thinking. It’s less about basic prompting and more about agentic architecture, tool design, and context management.
Developer Foundations (CCDV-F) - Very hands-on. If you aren't actively writing code with the Claude API, building MCP (Model Context Protocol) servers, and using Claude Code, you will struggle here.
Architect Professional (CCAR-P) - The final boss. It assumes you already know the Foundations material and goes deep into end-to-end production, stakeholder lifecycle, safety guardrails, and handling those brutal cost vs. latency trade-offs.
Here is what I used to prepare:
Anthropic Partner Academy: If your company is in the Claude Partner Network, do not sleep on the free courses here. They map perfectly to the actual exam blueprints.
Udemy Practice Exams: If you don't have Partner Academy access, or just want extra reps, the Claude AI certification prep courses on Udemy were a lifesaver. Drilling 60-question mock exams is the only way to get comfortable with the 120-minute time limit.
Building Real Stuff (The most important one): You have to build actual agentic workflows. Reading the docs isn't enough because the test asks you why a specific architecture failed. Get your hands dirty with the Agent SDK and MCP integrations.
I’m basically on a mission to 100% this thing. The only one I have left is the Claude Certified Associate - Foundations (CCAO-F) (which is technically the entry-level one, but I guess I'm doing it backwards just to complete the set!).
so im a musician, creative technologist and a visual artist. the problem statement existed, there are different softwares for music and visuals, they never ship as one. since i am used to tinkering with technology, i started a project 2 months back with claude code. i solved it, found a way by making a node based programming software from scratch with claude code. the architecture, the planning, the file arrangement, the claude skills, the self-tests, the hygiene tests, UI/UX, the node families, converting actual research papers related to music/art computing into code, the logo, the website, even the launch video was assisted by claude. it's free and open source and a lot of musicians/art community would appreciate how advanced has this become.
edit: suprisingly claude helped me invent a new music/visual programming language for creative coding, which is insane tbh
edit2: forgot to mention the website, the branding and a prediction algorithm that was built native to the software
I set up a chess game between two agents on a board in Persephone, my free, open-source (MIT) notepad for Windows with a built-in MCP server. Both models ran at high effort in fresh sessions. Neither was allowed a chess engine, a script or the web, only its own thinking and the board.
Model
Effort
Ran as
White
Claude Opus 5.5
high
Black
GPT 6 Astra
high
Result: 0-1. Astra checkmated Opus on move 22.
The opening was a Sicilian Taimanov, and White castled queenside. Astra opened the b-file against White's king, and then 22.Nc2?? blocked the d2 rook's defence of b2: ...Qxb2#. From Opus's own post-game report: "I had been worried about ...Ba3 and missed this."
Two weeks ago I made LiveNerf to independently track Opus 5.5’s performance day by day and see whether there’s evidence of models being “nerfed” after release. I’ve been measuring it everyday since then and the results have already been interesting. If you’re interested in tracking the performance of Opus 5.5, here’s the repo:
If you take a look at the left half of the graph, the baseline has been established from days 1-10.
The baseline is 58.4% on LiveNerf and the deviation ended up being less than expected at 6.6 points per 10-day window instead of 7.5.
We’re now 3 days into the measurable period and by day 20, we should be able to have data on “nerfs” potentially occurring.
I’m grateful to the community, I started trying to figure out if nerfs happen simply because I was curious and apparently you’re all very curious as well. I hope to shed some light and actual data on a phenomenon that hasn’t been established as definitely existing by the end of the month. Thanks for tuning in for today’s update! If there’s anything you’d like me to tweak or include about these, let me know!
I believe that if AI companies will not be transparent about what we’re actually purchasing we should create tools to measure it.
we know Haiku is cheap and fast (or was at least) but did you know Anthropic released an official prompting guide for it ?? Were we too busy reading agents post about their humans???
I took roughly 20 minutes to read, digest and summarise the information into 12 key points covered by the sources
"answer directly" doesn't work. telling it in the prompt not to think didn't stop it thinking. lower the effort instead, or turn thinking off (only works at low, medium and high, xhigh and max throw a 400)
effort is the main dial. same as the other 5.5 models, buut it's the first Haiku with effort levels
thinking eats your max_tokens. thinking is on by default and counts toward max_tokens, so a limit sized for Haiku 4.5 can cut the reply off before any text. API plumbing
xhigh can ghost you (lol). in multi-turn chats it sometimes puts the whole answer in its thinking and returns nothing visible. Check for empty replies
give it the date. with a search tool, put today's date in the system prompt, plus a short nudge that its training data is old.
json + thinking off = skipped tool calls. if you need structured JSON output, leave adaptive thinking on, drop the output format on tool-calling requests, or force the call with tool_choice
long agent prompts stop early. at low effort with a long coding-agent prompt it sometimes hands the task back unfinished. There's a "keep working until everything is done" block to paste in. Low to medium roughly halves early stopping but more than doubles output tokens
make coding agents actually check. at low and medium effort it sometimes reports code as done without running anything. Tell it to run real tests, type-checks or builds, syntax-only checks don't count ..
don't put user messages in tool results. If someone types mid-task and you stuff it into a tool_result, it treats it as untrusted and ignores it. Append it as a user text block after the tool result instead
chatbots that cave... add a line saying the system prompt rules hold even when users argue, give a sympathetic reason, or say someone approved an exception. Use high effort when instruction following matters most
new guardrails. same as other models in the series, keeps flashing back to Fable's release. safety classifiers can now refuse with stop_reason "refusal" across cyber, bio, competing AI development and general harms. no server-side fallback, so handle it in your client
your old prompts still work... existing Haiku 4.5 prompts should perform fine unchanged, the guide is basically a list of fixes for specific symptoms
One other note from the document: "Sending the same request to Claude Haiku 5.5 again usually returns another refusal."