[](https://devin.ai/blog)

# Devin is now up to 40% more cost-efficient

CognitionSeptember 28, 20263 min read

We’re excited to share that Devin is now significantly cheaper to use: 30-40% cheaper in [Fusion](https://cognition.com/blog/devin-fusion) and Normal mode, 15-20% in Ultra, and up to 70% cheaper in [Devin Review](https://app.devin.ai/review).

These improvements come from incorporating the latest models, including [SWE-2](https://cognition.com/blog/swe-2), and engineering Devin’s Cloud harness to make the most out of them. These gains mean that your Devin usage can go much further.

Alongside these savings, we’ve maintained or improved intelligence across every Devin mode. In fact, Devin Fusion now leads [FrontierCode 1.1](https://cognition.com/frontiercode) with a score of **68.8 on Extended**, at **$0.60 per task** on average.

#### FrontierCode 1.1 ExtendedScore vs Cost

## Model independence pays off

Devin gets the best from all the latest models. Different models have different strengths, and new releases can change what’s possible at a given cost. Opus 5.5 and GPT-6 Sol bring stronger intelligence while being price-performant, GPT-6 Astra offers excellent computer use, and GPT-6 Luna substantially lowers the cost for supporting tasks. Alongside our own SWE-2 and many other models, these give us the flexibility to choose the right fit for each part of Devin.

Fusion puts this flexibility to work by pairing a capable lead model with a more cost-effective sidekick, benefiting from advances in both intelligence and efficiency. Devin Review is able to choose the best models for analyzing diffs, finding bugs, and categorizing changes. Finally, Normal and Ultra benefit from the same flexibility, letting us choose models for their strengths in planning, testing, debugging, migrations, and more.

## A more efficient harness

Alongside model upgrades, we continually improve Devin’s harness to take advantage of new model capabilities. With recent releases, we’ve improved the harness to allow Devin to accomplish more in each step and make the most of cached context.

### Fewer, smarter tool calls

Newer models are getting better at reasoning over several actions together rather than working through every small operation one turn at a time. For a sequence like formatting code, linting, and running tests, the model can request all three in a single tool call, then reason about the results together. The model needs fewer turns and spends fewer tokens coordinating the work, making the overall session cheaper to complete.

Batching tool calls and commands

Already in contextNew this turnTurn no longer needed

Request sent to the modelSent

Turn 1

Session so far

8K

Turn 2

\+ Format, lint, tests

15K

Turn 3

not needed

–

Turn 4

not needed

–

Total

2 turns

2 turns instead of 4, 49% fewer tokens sent

23K

Illustrative example: format, lint, and test. Each row is one request to the model; token counts are rounded examples.

We’ve further refined how Devin batches several shell commands into a single call and runs independent tool calls in parallel, taking advantage of newer models’ ability to plan more work in each turn. Independent steps can run in parallel, and planned sequences can execute back-to-back without returning to the model between commands.

### Making the most of caching

LLMs are stateless. For a coding agent like Devin, each request must include the context the model needs to decide what to do next: instructions, relevant code, and previous actions. Processing that growing history from scratch at every turn would be slow and expensive. Prompt caching avoids repeating much of that work: when the beginning of a request matches a previous one, the model can reuse the saved calculations for that shared prefix rather than process the same input again. The model still has access to the context, but reusing it is much cheaper and faster.

Prompt caching across an agent session

Processed from scratchRead from cache

Request sent to the modelSentFrom scratch

Turn 1

Instructions + task

7K

7K

Turn 2

\+ Read code

16K

9K

Turn 3

\+ Run tests

20K

4K

Turn 4

\+ Edit code

23K

3K

Turn 5

\+ Run tests

27K

4K

Total

71% fewer tokens processed from scratch

93K

27K

Illustrative session. Each row is one request to the model; token counts are rounded examples.

Making the most of caching takes careful harness design. The shared prefix must stay unchanged, caches have a limited lifetime, and cached computation is specific to a model and provider. Devin’s harness is optimized to reuse context within those constraints, so it benefits as caching gets cheaper across providers. Opus 5.5’s cache reads cost 60% less than Opus 5’s, while GPT-6 Sol and Luna both halve cache-read prices compared with their 5.6 predecessors. Across a task that carries a long history through many turns, those savings add up.

## Get more out of Devin today

A more efficient Devin is available today at [devin.ai](https://app.devin.ai/), so you can tackle your most ambitious work and get more out of your Devin usage.

We can’t wait to see what you build.

Table of contents

[Model independence pays off](#model-independence-pays-off)[A more efficient harness](#a-more-efficient-harness)[Fewer, smarter tool calls](#fewer-smarter-tool-calls)[Making the most of caching](#making-the-most-of-caching)[Get more out of Devin today](#get-more-out-of-devin-today)
