Major Upgrade to the DinoAI Harness

We just shipped a major upgrade of the DinoAI harness with 88% lower cost, 6x lower token consumption, and in-built work verifiers.

Kaustav Mitra

·

·

7

min read

dino ai users see 10X less claude errors

In the world of data engineering, not many are attempting to automate it. A few that are, most of them would rather ship a shell and get you to bring your own token, build your own harness so they don't have to build the intelligence for actually getting data work done efficiently using AI agents. That's the problem we are solving for our customers. We want our customers to be able to go from POC to production, while most of our smaller competitors today are catering to the hobbyists.

When you build your own AI agent and direct it to solve data engineering problems, those agents will tell you they fixed even when the changes do not compile. This is because LLM models are uniquely biased to reach a positive result. As users, you find this out in CI, or later in production.

We have watched this happen enough times that we no longer treat it as a quirk. This happens because verification is optional when you are using a frontier model for agentic workflows.

For us the interesting question in agentic data work is can the agent platform ensure formal verification of the work is done when the work is completed. That will increase trust, accuracy, and enable users to extend to production use cases.

Today, we have customers using self-healing agents on top of their production data pipelines running on Bolt, or in their own cloud on Airflow or Dagster. For us, this is a problem that needs to be solved.

The first generation of the DinoAI harness we built, excelled at interactive co-pilot use cases. However, as more and more users started using DinoAI agents for background tasks, that harness started to fall short. It needed a fundamentally new design.

So, at Paradime, we completely rebuilt the DinoAI harness so that it works for both co-pilot and background agentic tasks. Three outcomes came out of the work, that we measured on live infrastructure:

  1. 120X spin up time improvement. Agent sessions now in almost a second, compared to 2 mins before.

  2. Up to 88% lower AI cost per task. Users saw the same work completed for a fraction of the spend.

  3. Runaway sessions eliminated. We have bounded context so any particular session will not go out of whack.

  4. Built in verifier. DinoAI has to prove its data work is correct before it is allowed to say "done." No other AI data agent enforces this out of the box.

Agent Spins up 60X faster

Paradime agents spin up was slow and almost always took somewhere between one and three minutes. We have now fixed that. We have completely re-architected our sandbox and now the sandbox spins up in less 2 seconds, almost instantly.

"Memory Sidecar" for agents

Running agents consume tokens and tokens are consumed at every step and turn. Any context that is added to the agent is carried through each turn. Now, if an user added a huge file or a giant search result in a turn, that context then would carry forward through each step and turn and the context size would keep monotonically increasing consuming tokens again and again.

So we fixed with the concept of "memory sidecar".

We park big artefacts on a memory layer that is attached as a sidecar to the agent session. It's like side notes that an analyst keeps while doing critical work. We fetch this information when we want and discard it when we don't need it anymore.

This keeps the context window range bound and prevents it from exploding. This also helps us manage unstructured context from Jira, Linear, Confluence, Notion, Google Drive, and more range bound and manageable.

Task type

Cost change

Reading large files

−72%

Processing big command outputs

−50 to −65%

Delegated sub-tasks

−53%

Editing large files

−81%, runaway sessions eliminated

One real runaway task used to burn about 600K tokens. It now finishes in about 95K. Same result, roughly 6× cheaper.

Reduced Context Rot

We also stopped stuffing the context with tool noise coming from vanilla MCP like wrappers. When DinoAI reads a Jira ticket, a Confluence page, or a Linear issue, those APIs return a lot of clutter. We now distill and remove that clutter at the source. For example, on Jira, we keep sprints, story points, owners, and dates while we discard the rest.

As a real-life measurement Jira ticket: 64,000 characters → 4,200. We reduced context rot by 93%, which keeps more of the useful fields. We also architected in a way that the full raw response is also available through a separate request if the agent needs it.

This is where a lot of "we connected MCP" anecdotes fall apart. If you bolt a vanilla MCP server onto a general coding agent, you typically get the raw API or tool payload dumped straight into context like the full Jira / Confluence / Linear JSON, comments, metadata, nested fields, the lot.

However, the agent and the task does not need that. Instead of pulling all that into context and then reasoning on it, it's faster and more accurate to pull what is needed. Context stays clean and outcomes are more predictable.

We built this into DinoAI where the context is filtered at source, before it's sent to the LLM models. Same Jira ticket, using DinoAI has 93% less characters in the working context.

Bolting MCP onto a general agent does not give you the same cost or quality. You get a wrapper that is bringing the whole API output like a firehose into your context. We instead chose to reduce the clutter.

When we measured this on a real task, the AI cost per task went down by a humungous 88%.

In-built Verification

Every AI agent on the market claims it's done the job that is not real. Most tools ask the model to double-check, build this into their system prompt and hope the model listens. We have seen with newer models like Opus 4.6+, Fable series, that these models are quite poor at following system prompts because their underlying system prompts have been laid so bar that any rules you give them, they are pretty bad at following.

So, we decided that was not good enough for production data engineering work including building dbt™ pipelines, so we stopped leaving verification to the model.

We have now built-in a verifier into the DinoAI harness so that the agent workflows can independently verify the files it changed. If the work is broken, the agent figures out the exact error, and fixes it. It cannot skip the verification.

The verification is adaptive and is tuned to the change:

  • Comment-only edit → quick syntax check. Your warehouse is never touched.

  • Real logic change → the warehouse validates the SQL in seconds, with near-zero cost.

  • Model with data-quality tests → build + tests run before DinoAI is allowed to complete.

In real-life examples, we told an agent to break a model and immediately claim success. The harness caught it, blocked the claim, and the agent fixed its own mistake. 25 seconds from bad edit to verified fix. In eval runs the gate also caught accidental mistakes twice and forced corrections before anyone saw them.

When something genuinely cannot be fixed, DinoAI says so. Retries are bounded, outcomes are transparent, and there are no silent failures or infinite loops. Broken code is caught, blocked, and fixed before it reaches the user.

Prompted checks ≠ enforced gate

A lot of agent products put validation tools next to the model e.g. compile, build, run tests, and call that reliability. This control model is problematic: if the agent chooses whether to call the check means this control is not deterministic.

We have made control deterministic in our agents.

We are shipping a pre-delivery gate that blocks "done" on unverified data work, with risk-based escalation so you never pay for warehouse checks when the change is as simple as a comment edit.

The verification layer is deterministic platform code, not another AI judging an AI. It is auditable, loggable, and fails safe.

No other AI data agent enforces this.

How we compare


Typical AI coding/data agents

DinoAI today

Verifying its own work

Optional. The AI is prompted to check and may skip it

Enforced by the platform. Cannot be skipped

Wrong answers reaching users

Whenever the model doesn't self-check

Broken data work is blocked before delivery

Cost of big files & searches

Carried in full, re-billed every step

Distilled and parked. Up to 88% cheaper

Runaway sessions

Known failure mode, surprise bills

Structurally bounded

Verification cost

All-or-nothing, if at all

Complexity is tuned to the change. Small changes are cheap to verify.

Honesty when stuck

Often silent or overconfident

Tells you what it couldn't resolve

The industry's own published results keep showing the same failure: even well-tuned agents still fail in the confidently-wrong way when verification is left to the model's discretion. We have made it a platform guarantee, instead of a prompt suggestion.

Provider by provider Comparison

Based on public repos, docs, and our own hands-on analysis as of September 2026.


Paradime DinoAI

Altimate (altimate-code)

OpenCode

Claude Code (vanilla)

Snowflake Cortex Code ("CoCo")

What it is

Managed, data-native autonomous agent (Slack, API, scheduled)

Data-focused CLI agent - an OpenCode fork with ~100 data tools

Open-source general coding agent (CLI)

General coding agent (CLI/cloud)

Snowflake's agentic coding assistant (Snowsight, CLI, desktop IDE) + Cortex Agents framework

Verifies work before claiming done

✅ Enforced by the platform — the agent cannot skip it

❌ Rich validation tools exist, but the AI chooses whether to call them

❌ Model's discretion

❌ Model's discretion (hooks let a power user hand-build checks)

🟡 Iterates on failures when it runs dbt™ - agentic habit, no enforced pre-delivery gate in their docs

dbt™-aware validation (compile / build / tests)

✅ Built-in, automatic, escalates with risk

🟡 Available as tools, model-invoked

❌

🟡 Only if the user prompts for it

🟡 Yes — for dbt™ Projects on Snowflake (scaffold, test, run, self-fix); dbt™ must live in their ecosystem

Cost engineering (context distillation, runaway bounds)

✅ Built-in; measured 50–88% less per task, runaways structurally eliminated

🟡 Model/prompt-dependent

🟡 Basic truncation

🟡 Good general hygiene; runaway cost still possible

🟡 Credit-billed by Snowflake; no published context-cost engineering or runaway bounds

Warehouse coverage

✅ Snowflake, BigQuery, Databricks, Redshift, DuckDB/MotherDuck…

✅ Multi

➖ Not warehouse-native

➖ Not warehouse-native

❌ Snowflake only

AI models & keys

✅ Managed by Paradime

❌ BYOK

❌ BYOK

🟡 Anthropic subscription/key

✅ Snowflake-hosted

Where the work lands

Your repo: real edits, branches, MRs/PRs — verified

Local files / git

Local files / git

Local files / git

dbt™ Projects on Snowflake workspaces (git-integrated), inside Snowsight

State-Aware-Orchestration in the Harness

Once the DinoAI agents finish a task, then the next step is running pipelines. Paradime State is how the harness decides what actually needs to run.

This week, we shipped Paradime State - state-aware orchestration for humans and DinoAI agents inside Paradime Bolt. Before a model is built, Paradime State is the Decision Engine on top of dbt Core™ v2.0, and chooses to SKIP, CLONE, or BUILD it from SQL logic, source freshness, object reuse across jobs, and column-level lineage. You only burn warehouse compute on what changed.

Combined with the harness, this is extremely powerful. Agents query and run pipelines differently compared to humans. So when an organization fires of a swarm of agents to fix all the tech debt and Jira tickets, having a cached state means the work will be faster and warehouse credits consumed will be much lower.

Paradime is now the only company in the world that supports state-aware orchestration in the harness - for humans and for agents.

You can read more and how it compares to state:modified+ here: State-aware orchestration for humans and agents.

What this means for your team

Data leaders: AI spend per task drops and becomes predictable. No runaway bills. Agent ships deterministically verified work and hence the agents can be rolled out into production.

Analytics engineers: DinoAI stops being a fast intern you have to babysit. Changes arrive pre-compiled and pre-tested, so the first pass is more often correct, and review burden goes down.

Conclusion

This update to the new generation of the Paradime harness results in radically faster spin up times, faster results using memory sidecar and accurate context, and finally in-built verification for accurate outcomes.

We will soon publish benchmark results we achieved with the new harness, but believe me when I say that this is gonna blow your mind. I have seen teams satisfied with having just a few skills in their repo and then how they got burned trying to wrangle LLMs with newer models. Eventually, they missed deadlines on getting AI into production.

On the other hand, with Paradime we are giving users the picks and shovels, the infra to build those agents and ship them into production.

If you want to know more, sign up for free at https://app.paradime.io and reach out to us.

Interested to Learn More?
Try Out the Free 14-Days Trial

More Articles

Stop Managing Pipelines. Start Shipping Them.

Join the teams that replaced manual dbt™ workflows with agentic AI. Free to start, no credit card required.

Stop Managing Pipelines. Start Shipping Them.

Join the teams that replaced manual dbt™ workflows with agentic AI. Free to start, no credit card required.

Stop Managing Pipelines. Start Shipping Them.

Join the teams that replaced manual dbt™ workflows with agentic AI. Free to start, no credit card required.

Copyright © 2026 Paradime Labs, Inc. Made with ❤️ in San Francisco ・ London

*dbt® and dbt Core® are federally registered trademarks of dbt Labs, Inc. in the United States and various jurisdictions around the world. Paradime is not a partner of dbt Labs. All rights therein are reserved to dbt Labs. Paradime is not a product or service of or endorsed by dbt Labs, Inc.

Copyright © 2026 Paradime Labs, Inc. Made with ❤️ in San Francisco ・ London

*dbt® and dbt Core® are federally registered trademarks of dbt Labs, Inc. in the United States and various jurisdictions around the world. Paradime is not a partner of dbt Labs. All rights therein are reserved to dbt Labs. Paradime is not a product or service of or endorsed by dbt Labs, Inc.

Copyright © 2026 Paradime Labs, Inc. Made with ❤️ in San Francisco ・ London

*dbt® and dbt Core® are federally registered trademarks of dbt Labs, Inc. in the United States and various jurisdictions around the world. Paradime is not a partner of dbt Labs. All rights therein are reserved to dbt Labs. Paradime is not a product or service of or endorsed by dbt Labs, Inc.