Claude Code After 18 Months: 20 Percent, Not 10x

• by Tobias Schäfer • 7 min read

I’ve been using Claude Code every day for about a year and a half now. Shopware plugins, Shopware migrations, a Laravel SaaS, this blog. If I’m honest, it saves me 20–25 percent of my time.

That’s a good number. It’s also nowhere near the 10x that lands in my timeline every week. And it’s a lot better than last summer’s headline about AI making you slower. Both come from somewhere, so I went back and read the studies properly before writing down my own numbers.

What the studies actually measured

The famous one is METR’s 2025 randomized trial. Sixteen experienced open-source developers worked on 246 issues from their own repos, mostly with Cursor and Claude 3.5/3.7 Sonnet. With AI, they took 19 percent longer. Beforehand they expected a 24 percent speedup. Afterwards they still thought they’d been 20 percent faster.

That gap between feeling and measurement is the most interesting part of the whole thing for me. Hold on to it for my numbers further down.

In February 2026 METR put out the follow-up, this time with late-2025 tools: 57 developers, 143 repos, over 800 tasks. The ten who had been in the first round now needed about 18 percent less time with AI (confidence interval −38 to +9 percent), the 47 new ones about 4 percent less (−15 to +9). Both intervals include zero. And the study itself started to slip. Between 30 and 50 percent of developers held tasks back because they didn’t want to do them without AI. Hard to randomize around that. METR calls its own result “only very weak evidence” for the size of the effect and is redesigning the study.

Then there’s the top end. METR went through 5,305 Claude Code transcripts from seven of its own people in January 2026 and got roughly 1.5x to 13x time savings on the tasks done with the agent. With caveats: they call it a soft upper bound themselves, and the LLM judge doing the estimating was only checked against 34 human labels. They also say the important bit out loud: “People likely do not create 10x as much value with AI, even if we observe a 10x time savings factor on tasks that people do with AI.”

In a May 2026 survey, 349 people in technical jobs said, at the median, they’re three times as fast and their work is worth 1.4–2 times as much. Right next to it, the same authors point out that in 2025 people overestimated AI’s effect on their time by 40 percentage points on average.

The big industry surveys fit the picture. According to DORA 2025, 90 percent of developers use AI, for a median of two hours a day. The report calls AI a “mirror and a multiplier”: it amplifies whatever the team already is. Rob Bowley looked at the report’s seven team clusters, and only two of them get more throughput without their change failure rate going up. In the Stack Overflow 2025 survey, 84 percent use AI tools or plan to, but 46 percent don’t trust what comes out. Two thirds are mostly annoyed by code that’s “almost right, but not quite”. 45 percent say debugging AI code takes longer.

So where things get measured properly, the result lands somewhere between slightly slower and maybe 20 percent faster. Where people estimate for themselves, it’s 2–3x. The 10x only exists as an upper bound on individual tasks. My 20–25 percent is an estimate too, by the way. Same discount applies.

Where it saves me time

Most clearly on small tasks in code I don’t know. A client shows up with a Shopware plugin someone else wrote three years ago. They need one more field in the admin, or a fix in an event subscriber. That used to be an hour, and most of it went into reading: Which event? Which service definition? Where is this thing even wired up? Now it’s about 15 minutes. The agent reads its way in, I look at where it ended up.

That fits the 2025 METR study better than I’d have guessed. Their developers were experts in repos they’d maintained for years. Which is exactly where I get the least out of it, too.

I use all the configuration there is: a CLAUDE.md per project, skills, hooks, MCP servers, subagents. Almost none of it comes from plugin marketplaces. What stuck is mostly stuff I built myself, because it works the way I work. One skill releases Shopware plugins with bilingual changelogs, another checks a plugin against the store rules before upload, and this post went through my own review commands. The off-the-shelf stuff was built for somebody else’s workflow. Bending it into shape took longer than writing my own.

Where it costs me time

Mostly when it makes things up. Claude Code is weirdly confident about Shopware APIs that don’t exist: repository methods and event names that sound exactly right. The nasty part is when you find out. Not locally, but in the test suite or on staging. So each one costs a full round trip instead of a second look. I feel for everyone in that Stack Overflow survey complaining about “almost right, but not quite”.

The other thing is runs that get out of hand. With some models, a two-line change has turned into a long session: half the repo read, tests run over and over, things rewritten that nobody asked for. You can’t really focus on anything else while that’s going. And sometimes the result is worse than the two lines you’d have typed yourself.

How I review, and why that eats into the win

In a codebase I know well, review is quick. Does it smell off? How big is the diff? Which files did it touch? If a ten-line fix changes four files, something’s wrong and I take a closer look.

In code I don’t know, I read until I actually understand the change. Otherwise I can’t tell plausible from correct. And that’s the uncomfortable bit: where the agent saves me the most time is exactly where I have to look hardest. Part of the 45 minutes I save on the unfamiliar plugin goes straight back into review. The 15 minutes only hold up if I don’t skip that.

The money

I pay €90 a month for the Max 5x plan. 20–25 percent of a full-time month is somewhere north of 30 hours. On paper, that’s not a hard call.

But remember those 40 percentage points. If I’m as far off as the 2025 average, I’m actually slower with the tool. I don’t think so. The tasks in unfamiliar codebases are the ones I can roughly time, and they’re clearly faster. For everything else, I can’t rule it out.

Am I getting worse at writing code?

Probably, yes. More at recalling things than at understanding them. When I’m stuck, I don’t dig around in my head anymore for stuff I used to know by heart. I just ask.

If it got switched off tomorrow, I’d need a few days to get used to it again. After that I’d hopefully be as productive as before. Although “hopefully” is carrying a lot of weight in that sentence.

So, 10x?

No. For me it’s a solid 20 percent, picked up mostly in unfamiliar code and partly handed back in review. Easily worth the €90 a month. I still can’t fully prove it.

If you want your own number rather than mine or METR’s, don’t ask yourself how much faster you feel. That’s exactly where people were 40 percentage points off in 2025. Take a few small tasks, write down beforehand how long each would take without the agent, and check against the clock afterwards. Crude, but better than gut feeling.

And if you get a very different number, in either direction: write to me and tell me what kind of projects you’re working on.

Related Posts