DeepSeek Harness Beats Claude Code On Speed (Tested)

Julian Goldie — founder, AI Profit Boardroom
By Julian Goldie · 9 min read
Get The AI Profit Stack Join AIPB →
🎯 1,000+ done-for-you AI agent workflows 📅 5 live coaching calls / week with me 🛡️ 7-day refund + 30-day ROI guarantee 👥 3,600+ AI operators inside

DeepSeek Harness built a full 3D animated website and a working game in 11 minutes.

Claude Code, on the identical prompt, was still building at 30 minutes.

That is the deepseek harness vs claude code headline, and it is real — we ran it, filmed it, and watched the token counters the whole way.

The catch is what each one produced at the end, and that is where most comparisons stop being useful.

Here is the full test.

The test in one paragraph

One prompt, no context, no follow-ups: build a 3D animated accountancy website and a simple game.

DeepSeek Harness ran DeepSeek V4 Pro.

Claude Code ran Claude Opus 5 on high.

Frontier model on both sides, so the comparison is about the harnesses and the models as shipped, not about a handicap.

Speed: the number everyone will quote

Eleven minutes, start to actual finish, for DeepSeek Harness.

Claude Code was past 30 minutes and still going.

Roughly three times the wall-clock time for the same brief.

If you ship a lot, that is the difference between iterating three times before lunch and iterating once.

Speed is not vanity when speed is attempts.

Quality: the number nobody quotes

Then we opened both.

The Claude site had animation tied to the mouse, so it felt responsive rather than decorated.

It also pulled in details it was never given, including a reference to Northwest England, because Claude had context from earlier sessions.

That is what "it knows my business" looks like in practice.

The DeepSeek site was animated and cartoony.

Its Tetris worked fine.

But a cartoon game on an accountancy website is a tonal mismatch, and it is the kind of thing a client spots before they read a single line of copy.

The Claude game also felt smoother.

Be clear about this though: neither build was finished.

Both needed more context and more prompting before they were worth publishing.

The full scoreboard

Metric DeepSeek Harness Claude Code
Wall-clock to finish 11 minutes 30+ minutes, still running
Model DeepSeek V4 Pro Claude Opus 5 (high)
Tokens burned 483,000 48,000 at 20 min
Price per token ~57x cheaper Baseline
Cost of this build ~5 cents Materially higher
Used context it was never given No Yes
Score 7 / 10 9 / 10
Stage v0.1 developer preview Mature product

The token story behind the speed

DeepSeek did not just finish faster.

It also burned 483,000 tokens doing it, against Claude's 48,000 at the twenty-minute mark.

Ten times the output for a less polished result.

That is what a very verbose model looks like when you actually measure it.

The saving grace is the price: about a 57th of Claude's cost per token, which is why the whole build came to around 5 cents.

The honest caveat is that this maths depends entirely on DeepSeek's current pricing.

Verbose plus cheap is a bargain.

Verbose plus normally-priced is not.

Why the harness itself matters more than the speed

DeepSeek Harness hit 105,000 GitHub stars in about two days, which puts it among the fastest growing open source projects ever.

That did not happen because of an 11-minute build.

It happened because the harness is free, open, and model-agnostic.

You run it locally, and you decide what brain goes in.

Want the cheapest possible setup? Plug a free model in — OpenCode works as a free brain inside it — and your running cost goes to zero.

That is a fundamentally different offer to a closed tool with a monthly fee and a fixed model.

And it is version 0.1.

A developer preview, roughly a tenth of what it will be at a proper release.

What I would do with this today

Claude Code keeps the work that gets seen.

DeepSeek Harness takes the volume: drafts, scaffolding, bulk generation, anything where the first three attempts are expected to be bad.

At 5 cents a run you stop rationing attempts, and that changes how you work more than any feature does.

I do not switch manually.

An orchestrator inside my agentic operating system picks the engine per job.

I never installed the harness myself either — I asked Claude to set it up, test it and wire it in, and that is the honest easiest route in.

🔥 Want the exact build?

Inside the AI Profit Boardroom I show the full Agent OS with DeepSeek Harness, Hermes and Claude Code side by side — plus weekly coaching calls and 4,000+ members.

→ Get access here

The verdict

DeepSeek Harness beats Claude Code on speed and on cost, decisively.

Claude Code beats DeepSeek Harness on the finished product, also decisively.

DeepSeek V4 Pro is a huge step up as a model, and the harness finally gives it a proper home.

But look at the two outputs and Claude still has a giant head start.

The right move is not to pick a winner.

It is to put both behind one orchestrator, and let the cheap one earn its keep on the work nobody sees.

FAQ

How much faster is DeepSeek Harness than Claude Code? Roughly three times on our test: 11 minutes to a finished build versus over 30 minutes and still running.

Which produced the better website? Claude Code. It looked more professional, the animation followed the mouse, and it used business context it was never given.

What did the DeepSeek build cost? About 5 cents, from a $10 account top-up, including a 3D animated site and a working game.

Is DeepSeek Harness free? The harness is free and open source. You supply the model, and free models work inside it.

Should I replace Claude Code with it? No — run both. Send high-volume, low-stakes work to DeepSeek Harness and keep client-facing builds on Claude.

Speed is really about attempts

The 11-minute number is easy to read as a bragging stat.

It is more useful than that.

An agent that finishes in 11 minutes gives you five attempts in an hour.

An agent that takes over 30 gives you one, maybe two.

Almost nothing good comes out of the first attempt, so the tool that lets you fail four more times in the same hour has a real advantage — as long as failing is cheap.

At around 5 cents a build, failing is effectively free.

That combination, fast plus cheap, is what makes DeepSeek Harness genuinely interesting even though Claude produced the better page.

You are not comparing one output to one output.

You are comparing one polished attempt to five rough ones, and for a lot of work five rough attempts is the better deal.

For client work it is not, which is exactly why both stay installed.

What the plugin design changes

The harness being open source is not a licensing detail.

It is the product.

You run the agent locally and you choose the brain that goes inside it — DeepSeek V4 Pro if you want the model it was built for, or a free model if you want the bill to disappear entirely.

OpenCode works as a free brain inside it, which is the route I would point a beginner at.

That means your setup is not hostage to one company's pricing page.

If the price moves, you move the engine.

If a better model lands next month, you plug it in.

That flexibility is worth more over a year than whichever tool is marginally ahead this week, and it is the reason 105,000 people starred the repo in two days.

A five-minute decision guide

If you want to stop reading and act, use this.

You ship client work and get paid for polish. Stay on Claude Code as your primary. Add the harness for drafts.

You generate at volume and cost is your ceiling. Make DeepSeek Harness your default and reserve the premium engine for finals.

You are just starting. Install the harness through an agent you already use, point it at a free brain, and spend nothing while you learn.

You already run an orchestrator. Wire both in and route per task. That is what I do, and it is the only version of this that scales.

You have ten minutes and no strong opinion. Run the same prompt through both, side by side, with no extra context. That single test will tell you more than any comparison post, including this one.

The honest limits of this test

We ran one prompt, once, with no context on either side.

That is a fair snapshot and it is not a full benchmark.

A different brief could flip the quality gap, and a longer session would let Claude's accumulated context matter even more than it did here.

What is solid is the shape of the result: DeepSeek Harness is much faster, much cheaper, much more verbose, and lands further from finished.

That shape has held every time I have used it since.

If you want certainty for your own work, run the same test yourself — one prompt, both agents, no extra context — and watch the token counters as well as the output.

It costs about 5 cents on the DeepSeek side to find out.

There are very few decisions in this business you can settle that cheaply.

About Julian

I'm Julian Goldie — AI entrepreneur, SEO expert, and founder of the AI Profit Boardroom (4,000+ members). I help business owners scale with AI agents, automation, and SEO.

→ Get my best AI training inside the AI Profit Boardroom

Related reading

On raw speed the deepseek harness vs claude code fight already has a winner — on finished work, it does not.

Real wins from inside the AI Profit Boardroom

See all 3,600+ members →
AIPB member win screenshot AIPB member win screenshot AIPB member win screenshot AIPB member win screenshot AIPB member win screenshot AIPB member win screenshot AIPB member win screenshot AIPB member win screenshot AIPB member win screenshot AIPB member win screenshot AIPB member win screenshot AIPB member win screenshot

Ready to Build AI Agents That Actually Make Money?

Join 3,600+ entrepreneurs inside the AI Profit Boardroom. Get 1,000+ plug-and-play AI agent workflows, daily coaching, and a community that holds you accountable.

Join The AI Agent Community →

7-Day No-Questions Refund • Cancel Anytime

← Back to all posts