Skip to main content
← All posts
|11 min read

Why not just use ChatGPT or Claude to run your business?

ChatGPT and Claude are great for asking questions and drafting one-off work. Running your back office every day is a different job: it needs direct API connections instead of built-in chat connectors, permissions and shared dashboards, and something that holds up on the odd customer, invoice, and document a chat window trips over.

The short version

ChatGPT and Claude are great for asking questions and drafting one-off work. Running your back office every day is a different job. It needs production-grade automation: direct API connections instead of the built-in chat connectors, permissions and shared dashboards so your team can rely on it, and something that holds up on the odd customer, invoice, and document a chat window trips over. That's the shape Aries is built in, and it covers the inference costs so a busy month never turns into a surprise bill.

Originally posted on the Aries blog page. Aries is built by Tundra AI Labs.

It's a fair question, and we get it a lot. You already pay for ChatGPT or Claude, they clearly understand your business when you explain it, and they'll happily draft an invoice reminder or summarize a work order. So why pay for anything else to run the admin side of your company?

A chat window and a system that runs your operations are two different tools that happen to share a language model underneath. One is brilliant at answering a question you bring it right now. The other has to do the same job correctly at 2 a.m. on a Sunday, across your whole team, without anyone sitting there to catch it when something goes wrong. Getting from the first to the second is most of the work, and it's the part that doesn't show up in a demo.

What ChatGPT and Claude are good at

Start with what these tools do well, because it's a lot. They're excellent at open-ended thinking: draft this email, explain this contract clause, rough out a pricing sheet, tell me what's wrong with this spreadsheet. You bring context, they reason over it, you take the answer and run. For that kind of one-shot help, a chat subscription is one of the best deals in software.

If your question is “can I use Claude or ChatGPT to help me work faster,” the answer is an easy yes, and you should. The trouble starts when the question becomes “can I use them to run a repeating part of my business without me in the loop.” That's a much harder problem, and it's not the one a chat is built to solve.

The gap between a good answer and a job that runs itself

Ask a chat to close out a finished work order and invoice it in QuickBooks, and it can usually do it, either by walking you through the steps or, with a connector wired up, taking care of it. Ask it to do that same job forty times a week, every week, while you're out on a site, and it gets less dependable. It handles the standard job fine and stalls on the one with a partial payment and a change order. You end up re-prompting, double-checking, and pasting things back and forth, which is fine once and exhausting as a daily routine.

A chat is built to be steered. Every good result assumes a person is there to phrase the request, judge the output, and nudge it when it drifts. Take the person away and you learn how much of the reliability was coming from them. Back-office work is the opposite of a one-off: it's the same handful of tasks repeated week after week, and it has to land right every single time, because a missed invoice or a dropped customer text costs you money.

Vibe-coded automations break on the edge cases

The natural next move is to build your way out of it. Wire the chat up to an API key, stitch together a few automation tools, maybe have the model write the glue code. It works in the demo, everyone's impressed, and for about a week it feels like you cracked it.

Then the exceptions start arriving. A customer's name has an accent in it and the matching logic chokes. An invoice comes in with the total in a different spot than the last ten. A supplier changes their PDF layout. The API you're calling returns a shape the script never expected. None of these are exotic, they turn up most weeks, and a solution that was assembled quickly, without anyone thinking hard about what happens on the unhappy path, tends to fall over on exactly these cases without any warning, so you don't find out until a customer does.

Software that holds up in day-to-day operations is mostly edge-case handling. What happens when the connection drops halfway through, when the input is malformed, when two things try to update the same record, when a run needs to pick up where it left off. That unglamorous work is the difference between a script that impresses on Friday and a system you can still trust in March. Vibe-coded tooling skips it, because skipping it is what made it fast.

Someone has to pay for the inference

A $20 chat subscription will happily run connector and MCP tasks for you, and for light use that's a great deal. The trouble shows up when you lean on it for automated volume, reading every invoice, drafting every reply, checking every work order. That's when you start running into limits: usage caps, rate limits, and throttling that slow the model down or cut it off right when the work is busiest. To push past them you either move up to pricier tiers or build on a raw API key, and once you're on the API every call is metered and billed straight to you, so the cost climbs with your volume.

Plenty of people build something on their own key that works in a demo, roll it out, and find a month later that the limits and the bill both scale with usage in a direction they didn't plan for. It doesn't make the idea wrong, it just means the cost of running it yourself climbs with how hard you use it, which is worth knowing before you commit. With Aries the inference is included in the plan, so heavy weeks don't hit a usage wall or turn into a surprise bill, and you can let it work as hard as the job needs.

One person's chat isn't a system your team can share

A chat session lives with one person. The prompts you tuned, the context you fed it, the little dashboard you asked it to render, all of that is trapped in your window and your history. Your office manager can't open the same setup tomorrow, your bookkeeper can't see the same numbers, and when you're out, the knowledge is out with you. To share it, everyone re-explains the business from scratch and hopes they phrase it the same way you did.

Running a company on that is fragile. Operations need workflows and reports that persist, that the whole team sees the same way, and that don't vanish when a chat history gets cleared. They need permissions, so the front desk can trigger the safe stuff while anything touching the books waits for a sign-off. They need an audit trail, so you can see what got done and who approved it. A chat product isn't trying to give you any of that, because that's not what it's for. It's a personal assistant, not shared infrastructure.

Built-in connectors time out; APIs hold

Here's the one that catches people off guard. To act on your tools, a chat app reaches them through built-in connectors, more and more of them built on MCP, the protocol these apps use to plug into outside services. Those connectors are designed for a live, interactive session. They have short timeouts, tighter limits, and a narrow slice of what the underlying tool can actually do, and they tend to drop the moment a job runs long or a session ends. That's fine for pulling up an order while you're looking at it, and rough for reconciling two hundred records overnight while nobody's watching.

A direct API integration works differently. It's the server-to-server connection the tool's makers built for exactly this, automated calls at volume, running unattended, retrying cleanly when something hiccups, reaching the full set of actions the software supports rather than the handful a chat connector exposes. It stays up. It doesn't need a human session behind it. When Aries connects to QuickBooks, MaintainX, Zendesk, or your email, it goes through those official APIs, which is why it can do far more than the built-in connectors and hold steady under repeated, everyday use instead of timing out on you at the worst moment.

What production-grade actually buys you

Put those pieces together and you get the thing a chat window can't be on its own: an automation that runs the same way every day, connects to the systems you already pay for through their official APIs, keeps your team on the same page with shared workflows and reports, asks for approval on the calls you want eyes on, and keeps a record of everything it touched.

Aries works this way. It sits on top of the tools you already run, handles the repetitive admin, and shows you the source next to its work so someone can approve it before anything changes. We go deeper on building automation around a review step in our guide to AI internal tools for small business.

None of this is magic the chat models don't have. It's the same models, plus the engineering around them that turns a clever answer into a job you can hand off and stop thinking about. Getting that engineering right, and keeping it running as your tools change underneath it, is the actual work, and it's the part that doesn't fit in a chat window.

When ChatGPT or Claude is the right call

Plenty of the time the chat is the right answer, and you shouldn't over-build. If the task is a one-off, if a person is going to look at the result anyway, or if you're still figuring out whether a workflow is worth automating at all, open ChatGPT or Claude and get your answer. They're the fastest way to think out loud and draft something you'll review yourself.

The line is repetition and trust. The day a task turns into “this same thing, over and over, and it needs to be right without me watching,” you've crossed out of what a chat is good at and into what a production system is for. Most back offices are full of the second kind of work, which is why “just use ChatGPT” takes you further than you'd expect and then stops well short of what the work actually needs.

Frequently asked questions

Can I use ChatGPT or Claude to run my business operations?

For one-off help, absolutely. For repeating back-office work that has to run correctly without you watching, a chat window falls short: it needs a person steering it, it forgets steps across sessions, it can't share workflows or permissions across your team, and its built-in connectors aren't built for unattended, high-volume automation, which is the job a production system is there to do.

Why do homemade AI automations get expensive?

A $20 chat plan is fine for light use, but lean on it for automated volume and you start hitting usage caps and rate limits. Getting past those means pricier tiers or building on a raw API key, where every call is metered and billed to you, so the cost rises with your activity. Aries includes the inference in the plan, so heavy usage doesn't hit a wall or turn into a surprise bill.

Why do the built-in AI connectors time out?

Chat connectors, including MCP-based ones, are designed for short interactive sessions, with timeouts and tight limits, and they drop when a job runs long. A direct server-to-server API integration is built for automated, repeated calls, stays connected, retries on failures, and reaches the full set of actions a tool supports.

Is a language model still doing the work in a tool like Aries?

Yes. The same class of model handles the language and judgment. The difference is everything around it: direct API connections, edge-case handling, shared workflows and reports, permissions, approvals, and an audit trail. All of that is what turns a good answer into a job you can hand off.

Putting AI to work in your business?

Talk to Tundra →