You're Building Workflows When You Should Be Building a Harness
On this page
- Three years ago, nobody would have believed this
- Most companies aren't using what's already there
- What it should look like
- What an agent harness actually is
- The five primitives I can name off the top of my head
- Why customers haven't caught on yet
- Why you can't generalize
- This is your only moat, and it won't last
- "But we have proprietary data"
- Case study: Harvey vs Claude Code
- Why the labs haven't come for you yet
- The objections I usually hear
- What happens next
Three years ago, nobody would have believed this
If I had gone back three years and told anyone that an AI agent could automate a real workflow in accounting, coding, or logistics, they would have laughed at me. Today it's so normal that nobody even talks about it anymore.
What I find ironic is that most applied AI companies were started in that same window, and they were built to make an unreliable AI useful. The models got reliable, but a lot of these companies still work like the models are flaky, and they're carrying habits from a time that's already over.
Most companies aren't using what's already there
I've worked with a lot of these companies, and most of them use only a small part of what a modern agent harness can do. Their engineers spend their time keeping fragile AI flows from breaking, learning how each customer works, and then building another fragile flow because the last one sort of held up.
Let me tell you what this looked like at the companies I worked at. An FDE would sit with a customer to learn their workflow and then write up action items for an engineer. The engineer would build the workflow and hand it back to the FDE, who would take it to the customer. Information got lost at every handoff, so what we shipped was brittle, and as soon as the customer wanted changes, we went through the whole loop again.
It took three weeks to ship one workflow for one customer, and we couldn't reuse it for the next one.
The part that honestly drives me crazy is that the engineers were already using coding agents to build and change these workflows. The agent was doing the actual work, and we put three people and three weeks between it and the customer, even though nobody understands a customer's workflow better than the customer does.
What it should look like
Take a boring, real workflow: every time a new email comes in, pull the key terms out of the attached contract and post them to the CRM.
Most applied AI companies would build this with a trigger node, a custom email integration, a custom CRM integration, a parsing pipeline, and a pile of glue code, tests, and dashboards. That's weeks of work, and it breaks the first time the contract format changes.
The other way is to give Claude Code access to the email inbox and the CRM and ask for it in chat. It just does it. There's nothing to build, and this works today, not in some future demo. When the customer wants something changed, you edit a sentence instead of rebuilding a pipeline.
What an agent harness actually is
An agent harness is everything around the model that lets it do real work: its tools, its context, its permissions, and its memory. Claude Code and Codex are both harnesses.
Engineers already work this way every day. They set up skills, APIs, and codebase context so the agent can take on almost any task they throw at it. If you give other industries the same setup, I think most AI automation gets solved.
The problem is that these coding tools are still too raw for a lawyer, an accountant, or an operations manager to use on their own, and nobody has figured out how to package them for everyone yet. That gap is the real opportunity.
The five primitives I can name off the top of my head
When I say "serve a harness properly," this is what I mean:
- Auth for every outside platform your customer uses
- Observability, so you can see what the agent is doing right now
- Traceability, so there's an audit trail of what it did and why
- Skills, which capture how this specific customer does this specific work
- Memory that carries over across tasks, people, and time
Once these are in place, customizing for a customer stops being a two-week sprint and becomes editing a skill file. Your FDE stops writing tickets for engineers and writes the skill with the customer, on the same day.
It really annoys me how much applied AI companies overcomplicate this, because the answer is simple. Serve a harness with these five things built in properly.
Why customers haven't caught on yet
Companies are getting away with the old approach right now, and it isn't because they're doing it well. There are three reasons.
First, most customers don't know these tools exist. That's the biggest advantage in the market right now, and it won't last.
Second, the customers who do know can't share what they build. Someone on the team makes a one-off automation in Claude Code, and it lives and dies on their laptop because there's no way to hand it to anyone else.
Third, observability and audit trails are missing from the raw tools. Enterprises need to see what their agents did, and that's the part they're actually paying you for, whether you realize it or not.
Why you can't generalize
Applied AI companies struggle to generalize, and their usual fix is to hire more people. In my experience it comes down to custom automations for each customer, no real process for turning FDE requests into product, and moving too fast without solid fundamentals in the codebase.
All three come from the same root problem, which is that customization lives in code instead of configuration. A harness fixes that, and more hiring doesn't.
This is your only moat, and it won't last
Integrating these primitives well is the one real advantage applied AI companies have over the big labs right now.
If the labs figure out how to serve it themselves, they solve their distribution problem at the same time, and they're already heading that way.
"But we have proprietary data"
I don't buy it. Most of the workflow data you're sitting on belongs to your customer, not to you. And if your vertical is big enough to matter, a lab can license or buy the data it needs to improve the general model, which keeps getting better faster than your fine-tune does.
Case study: Harvey vs Claude Code
You don't have to take my word for it. Go read r/biglaw, where a recent thread is titled, word for word:
Harvey is so bad compared to Claude Code it's infuriating.
What caught my attention is what the lawyers are actually complaining about, because it's not the model. Harvey lets you pick Claude Opus 5, the same model you'd get by going to Claude directly. Every complaint is about the layer Harvey puts around it:
- You can pick the model but not the effort level, so users suspect their requests run at lower effort than they'd choose. If you use Claude directly, you can pick anything from low to max.
- People feel the model is weaker inside Harvey than outside it, and they blame Harvey's system prompt.
- The features are a black box. The person who started the thread admits they have no idea what Harvey's "deep analysis" mode does, and another commenter didn't know Harvey could override the model they picked.
- It loses context in the middle of a conversation, and the original poster says the whole experience is worse than plain Claude.
- Most users, juniors included, don't know they can switch models, let alone why they'd want to.
My favorite comment pointed out that Harvey's CEO credited the company's revenue growth entirely to the product, while its forward-deployed legal engineers are out writing prompts for every practice group they can get in front of.
This is the exact pattern I've been describing: the same model with a worse harness, and people filling in the gaps by hand.
The pricing comparisons point the same way. For writing and analysis, general models are about as good or arguably better at roughly $25 a month per user, compared with $1,200 or more per seat for Harvey. Where Harvey still wins is its agent builder, Westlaw citation grounding, high-volume document processing, and firm-wide security, which are all primitives. And Harvey runs on models from the same labs it competes with.
So Harvey's advantage comes from integration rather than intelligence, and the people who use it every day are saying that integration is getting in their way.
To be fair, the same thread has some pushback. One commenter says the nerfing happens on Anthropic's side too and that Claude Code has the same problem. Another says a partner burned through $50k of Claude in one weekend. Going direct isn't automatically cheaper at heavy usage, and the labs have their own harness problems. I don't think that weakens my point, though. It means whoever serves these primitives best will win, and right now nobody is doing it well.
Why the labs haven't come for you yet
The labs aren't in your vertical yet for one reason: there isn't enough money in it for them yet. For that to change, more people need to understand how to use these tools, the market needs a clearer sense of what can and can't be automated, and the labs need to do their own customer discovery.
On top of that, OpenAI and Anthropic are locked in a fight over coding, and if either one slips there, its revenue takes a serious hit. That's where their attention is right now, and it's the only reason you still have time.
The objections I usually hear
"Enterprises need security and compliance." They do, and that's auth, traceability, and audit logs, which are primitives. You build them once, not once per customer.
"Our FDEs and customer relationships are our moat." Relationships matter, but an FDE who writes tickets for engineers is a bottleneck, while an FDE who writes skills directly with the customer gives you leverage.
"The agent will make mistakes in production." So will your brittle workflow. The difference is that with observability and traces you can see why it failed and fix it in a skill, instead of waiting on a two-week sprint.
What happens next
If the labs place a few bets and figure out how to integrate these primitives properly, I don't see what stops them from taking the entire enterprise market.
If you're still shipping one custom workflow per customer every three weeks, you're building a backlog instead of a company, and I think you're going to run out of time a lot sooner than you expect.