All articles
Real-world local AI experiment18 min read

A 50-Token/s AI Coding Agent for $0.08/Hour

Qwen3.8-27B · 1× RTX 3090 · ~96K context · no per-token billing

From failed fine-tuning experiments to Ollama, GPU rentals, DeepSeek Harness, OpenCode, and a one-command deployment setup.

MK

Mohammad Khosrotabar

Full-Stack Developer

$0.08

per GPU hour

40–50

generated tok/s

~96K

configured context

24 GB

RTX 3090 VRAM

At my observed rental rate: ~$12.80 for 160 GPU hours

The question

How far can an inexpensive self-hosted coding model actually go?

A few years ago, running a local or open model usually meant accepting a large quality or speed compromise. That is changing.

I rented a single RTX 3090 for approximately $0.08 per hour and ran Qwen3.8-27B Q4 with roughly a 96K configured context window. In favorable coding and chat workloads, I observed around 40–50 generated tokens per second. This was not a theoretical benchmark—it was a genuinely useful coding-agent setup on a consumer GPU.

8 hours/day

20 working days

160 hours total

$0.08/hour

$12.80

At that rate, 160 working hours of GPU access would cost approximately $12.80. That is GPU rental cost, not an unlimited AI subscription: marketplace prices vary, the server costs money while it runs, and context and model limits still exist. I stop or destroy the machine when I finish. While it is running, however, there is no additional per-token API bill.

I was surprised that an eight-cent-per-hour RTX 3090 was already this useful.

That result led to the more interesting experiment behind this article. I tested Java, Go, backend work, React, Next.js, frontend interfaces, existing repositories, greenfield applications, coding agents, and ordinary chat and reasoning. I tried fine-tuning, multiple GPU configurations, Ollama, DeepSeek Harness, OpenCode, and disposable rentals on Clore and Vast. I also used the model on work connected to a large Java banking project; my Java team lead was satisfied with the result.

My observed result, not a guaranteed RTX 3090 benchmark. Active context, prompt length, inference runtime, host quality and power limits, CPU and RAM, model version, caching, and the agent workload can all change performance.

The larger story became a progression: first, how cheap can a capable setup get? Then, what changes when I add another 3090, move to a 5090, or pay for enough VRAM to run BF16? The answer was more nuanced than “local wins” or “frontier wins,” and it showed me that inference software, context strategy, and the coding harness shape the experience almost as much as the model.


Model choice

Why a 27B model is an interesting middle ground

Very small models are cheap to run, but I often feel the loss in planning, multi-file coding, and recovery from tool errors. Huge models bring stronger capability, but their memory and rental requirements rise quickly. A dense model around 27 billion parameters sits in a useful middle: large enough for meaningful reasoning and coding, but small enough to fit on accessible consumer hardware once quantized.

Qwen3.8-27B also targets the work I care about: coding, tool use, agent execution, long context, frontend generation, and backend implementation. Its official model card specifies a native 262,144-token context window. A specification is not the same as a recommendation, though; later I will explain why I usually configured less.

Published evidence

The benchmarks are surprising—with an asterisk

The table below reproduces selected figures published by Qwen for Qwen3.8-27B and the Opus 4.6 Max comparison column. These results put a compact model unexpectedly close to—or ahead of—a much larger commercial system on several coding and agent tasks.

BenchmarkQwen3.8-27BOpus 4.6 MaxMeasures
Terminal Bench 2.173.078.2Agentic terminal coding
SWE-bench Pro61.753.4Agentic software engineering
NL2Repo-Bench42.347.6Repository-level generation
QwenSWEBench79.063.8Software engineering
CoWorkBench70.768.2Long-horizon work
LiveCodeBench v690.388.8Competitive coding

Source: the official Qwen3.8-27B model card, accessed September 2026. Qwen notes that evaluation harnesses and settings differ by benchmark; for example, its SWE-bench Pro comparison uses a Claude Code harness for most evaluated models.

Read this as vendor-published evidence, not independent proof. Benchmarks are sensitive to prompts, harnesses, graders, and task corrections. They explain why I took the model seriously; they do not prove it replaces a frontier coding agent in a messy repository.

Precision

BF16 versus Q4: the hardware economics changed completely

I tested both the full BF16 model, approximately 56 GB, and a Q4_K_M quantized build around 17–18 GB. BF16 keeps much more numerical precision. Q4 compresses the model’s weights into roughly four bits per value, with some values handled more carefully than others. Think of it as storing a high-resolution map with fewer shades: it is smaller and faster to move, but it is not mathematically identical.

Q4_K_M

~17–18 GB

Fits accessible consumer GPUs, offers much better price/performance, and became my daily default.

BF16

~56 GB

Preserves more precision, but requires high-VRAM or multi-GPU hardware and raises rental cost.

Q4 inevitably discards information. Yet in my practical coding workload it remained surprisingly capable, and the lower hardware requirement changed the entire cost equation. Most of the everyday results in this article came from Q4.

The failed path

My fine-tune worked—and still was the wrong answer

Customize first

I started with Unsloth, LoRA, a custom dataset, and multiple training epochs.

Compatibility work

I moved through inference settings, serving frameworks, quantization approaches, and CUDA, Triton, and runtime mismatches.

Technically successful

The adapter trained and ran, but the practical improvement did not justify the engineering complexity. At one point inference was only about 8 tokens per second.

Do not fine-tune until you have proved that the base model, prompting, context, runtime, agent harness, and tools are insufficient.

For my use case, inference and tooling optimization created far more value than fine-tuning. Fine-tuning is useful when you can name the behavior you need, measure it, and show that cheaper interventions fail. I had started with the most complex lever before validating the simpler ones.

The turning point

Ollama made the system boring—in the best way

Ollama is a local model runner and API server. It takes care of downloading a model, loading it onto the GPU, applying context configuration, running inference, and exposing an HTTP interface. That removed an entire layer of fragile setup.

Its OpenAI-compatible API was especially useful. Coding tools already understand that common request shape, so the rest of the system could treat my rented GPU like another model provider.

OpenCode / DeepSeek Harness / Chat
OpenAI-compatible API
Ollama
Qwen3.8-27B
GPU

Automation

From a fresh server to a coding endpoint in one command

I packaged the working setup into an open-source repository called local-ai-coding-box. The goal is simple: rent a fresh Ubuntu GPU server, clone one repository, run one command, and get a working coding server.

Recommended Q4 setup

./bootstrap.sh 128 quant

Full BF16 at 128K

./bootstrap.sh 128 full

256K examples

./bootstrap.sh 256 full
./bootstrap.sh 256 quant

The first argument selects the context target: 64, 128, 256, or auto. The second selects precision. quant maps to the roughly 17–18 GB qwen3.8:27b Q4_K_M model; full maps to the roughly 56 GB qwen3.8:27b-bf16 model. Quant is the default, so ./bootstrap.sh 128 is equivalent to ./bootstrap.sh 128 quant.

The script detects Ubuntu and the NVIDIA GPU, installs required packages and Ollama when necessary, writes configuration, applies the context choice, selects the model variant, skips an already-downloaded model, starts Ollama persistently, and performs health checks.

The systemd assumption that broke on cheap containers

My first version assumed every server had systemd. On Clore container hosts it hit a blunt error: systemctl: command not found. Many inexpensive GPU rentals are containers, not full virtual machines, so they do not run the usual service manager.

The fix: if systemd exists, use it. Otherwise, run Ollama inside a persistent tmux session.

That persistence matters because an SSH disconnect should not terminate the inference server halfway through an agent task.

Observed workloads

What I saw across 3090, 5090, and RTX PRO hardware

These are real observations from my workloads, not controlled laboratory benchmarks. Prompt length, occupied context, model variant, GPU topology, CPU, RAM, provider, caching, thermals, and software version all change the result.

My observed results

RTX 3090: the surprisingly cheap sweet spot

At the rental prices I found, the older RTX 3090 delivered the strongest budget story in my tests. These prices are snapshots from a marketplace, not fixed rates, but they show how inexpensive sustained coding inference can become when the model fits comfortably in consumer VRAM.

Budget configuration

1× RTX 3090

24 GB VRAM · Qwen3.8-27B Q4

$0.08per hour

~96K

Configured context

~40–50

Generated tok/s

160 hours × $0.08

$12.80 / month

More headroom

2× RTX 3090

48 GB aggregate VRAM · Qwen3.8-27B Q4

$0.18per hour

Larger

Context headroom

~80

Generated tok/s

160 hours × $0.18

$28.80 / month

The practical takeaway: for roughly $12.80 in GPU rental cost over 160 working hours, my single-3090 setup produced around 40–50 tok/s with approximately 96K configured context. At roughly $28.80 for the same 160 hours, my dual-3090 setup reached around 80 tok/s in one favorable configuration and gave me more context headroom.

This is not unlimited compute. I had to stop or destroy the machine when I was finished to avoid paying for idle time.

Two RTX 3090s do not universally mean twice the speed. Multi-GPU results depend on GPU communication, model distribution, inference runtime, context size, prompt processing, host CPU and RAM, PCIe or NVLink topology, and current context utilization. The approximately 80 tok/s figure is what I observed in my setup, not a guarantee for every dual-3090 host.

Configured context is also different from occupied context. A model configured for 96K or 128K can remain fast while only a small part of that window is filled. As an agent session accumulates files, logs, and conversation history, throughput can fall significantly.

HardwareVRAMModelContext workWhat I observed
1× RTX 309024 GBQ4~96K (also tested 128K)~40–50 tok/s in favorable everyday workloads
2× RTX 309048 GB aggregateQ4Larger context headroom~80 tok/s in one favorable real-world configuration
1× RTX 509032 GBQ4128KUp to ~120 tok/s in favorable workloads
2× RTX 509064 GB aggregateQ4Large-context agent workUsed for a substantial full-app first implementation
RTX PRO / high VRAMVariesBF16Large model/context headroomBest fit for full precision; considerably pricier

On one RTX 3090, Q4 at around 96K configured context commonly produced roughly 40–50 generated tokens per second in favorable normal coding and chat work. I also tried 128K. During a very long harness session, once the context was heavily filled, throughput eventually fell to around 8 tokens per second.

What if I want more speed?

A single RTX 5090 rental cost me about $0.26 per hour. With 128K configured context, I observed peaks around 120 tokens per second in favorable workloads. Two 3090s brought 48 GB of aggregate VRAM and useful 128K/256K headroom, but a second GPU does not automatically double decoding speed: data must cross devices and synchronization has a cost. Two 5090s gave me far more room for autonomous work. High-VRAM RTX PRO hardware was the practical path to BF16, at a much higher rental price.

Memory pressure

Configured context is not occupied context

A model supporting 256K does not mean 256K is always the best setting. Larger windows reserve or permit more KV cache—the working memory the model uses to track previous tokens. They can consume more VRAM, make long prompts more expensive to process, and reduce perceived responsiveness.

Configured context

The maximum window available to the session.

Occupied context

The portion currently filled by prompts, files, tool results, and responses.

A model configured for 128K but using only 8K can still feel fast. The same model near the top of that window has much more history to process and store. My practical sweet spot was usually 96K–128K: enough room for serious coding without paying the full memory and latency cost of 256K from the start.

The headline math

Where “about $42 per month” comes from

8 hours/day×20 days×$0.26/hour=$41.60

160 rented GPU hours in a working month, at one price I personally found.

There is no per-token API charge while I rent that machine. This is not literally unlimited compute. GPU time is finite, context and output limits still exist, and I pay while the server sits idle. The useful distinction is billing: a cloud API generally charges by tokens or requests; a self-hosted rented GPU charges primarily by time.

Destroy, rent, repeat

When I finish work, I destroy the server so it stops costing money. The next day I may rent a completely different host. Reinstalling the stack manually each time is tedious—that disposable-server workflow is exactly why the bootstrap repository became valuable.

Rental notes

Clore and Vast: two practical connection patterns

Clore.ai: mapped public ports

When creating the container, expose container port 11434 over TCP. Ollama listens internally at localhost:11434, while Clore may publish it on a different host and port such as PUBLIC_HOST:2787. An external coding tool would then point to http://PUBLIC_HOST:PUBLIC_PORT/v1.

Verify on the server

curl http://localhost:11434/v1/models
ollama ps

The first command checks the OpenAI-compatible models endpoint. The second shows which model Ollama currently has loaded. Clore can be attractive because consumer GPUs are often inexpensive, but do not casually expose an unauthenticated model server to the public internet. Native Ollama does not provide the same API-key security model as a commercial provider. Prefer an SSH tunnel, provider firewall, VPN, or authenticated reverse proxy.

Vast.ai: tunnel the remote port to localhost

Run from your computer

ssh -p SSH_PORT root@SERVER_HOST \
  -L 11434:localhost:11434

Then test locally

curl http://localhost:11434/v1/models

OpenCode or DeepSeek Harness can now use http://localhost:11434/v1. Inference still happens on the remote GPU. “localhost” is simply your end of the encrypted SSH tunnel, which is both convenient and safer than a bare public port.

Real work

Two application tests mattered more than synthetic prompts

A portfolio design

I gave Qwen a portfolio website task and it produced the site quickly. For this particular design task, I personally preferred Qwen’s visual result over some versions I had generated with ChatGPT and Claude. That is subjective, not a model ranking, but it made UI generation one of the strongest use cases I found—especially React, Next.js, Tailwind-style interfaces, landing pages, and dashboards.

A personal accounting application with flutter

With Qwen3.8-27B and one RTX 5090s, I asked for a interface, local-first storage, backup architecture, onboarding, accounting features, and a real application structure. A substantial first implementation emerged in about 80 minutes. The GPU rental cost for that session was approximately $0.80.

That does not mean production-ready in 80 minutes. It means a major working first version existed. Security review, domain validation, testing, accessibility, and ordinary engineering judgment were still required.

The honest split

Where Qwen impressed me—and where Codex is still better

Strongest experiences

  • Greenfield applications and prototypes
  • Frontend UI, React, Next.js, landing pages, and dashboards
  • Repetitive implementation and many backend tasks
  • High-volume developer chat

I still choose Codex for

  • Deep architecture and complicated refactors
  • Difficult debugging
  • Large, established repositories
  • Long-horizon software engineering

Starting from an empty directory, Qwen can move extremely quickly. In a mature repository it may spend a long time reading files, reconstructing architecture, tracing dependencies, and learning why earlier decisions exist. Sometimes that discovery phase approaches the time required to generate a smaller greenfield app.

Consider a Keycloak-style authorization system spread across Go services: resources, scopes, roles, policies, permissions, synchronization, and existing business rules all interact. I believe Qwen could work through it. Today I would still trust Codex more with that depth of architecture and refactoring.

My workflow is not Qwen instead of Codex. It is Qwen plus Codex: cheap volume for one, difficult high-leverage engineering for the other.

Agent tooling

The harness matters almost as much as the model

A coding harness is the working environment around the model. It gives the model repository access, file reading and editing, terminal commands, search, multi-step behavior, persistent sessions, and context management. A strong model without those capabilities is still mostly a chatbot.

DeepSeek Harness

I connected DeepSeek Harness to Ollama through an OpenAI-compatible endpoint. For the Q4 setup, the conceptual configuration is a base URL such as http://SERVER:PORT/v1 and the exact model identifier qwen3.8:27b. Through an SSH tunnel, the base URL becomes http://localhost:11434/v1.

The exact identifier matters: the provider and Ollama must agree on the model name. I liked Harness for autonomous workflows, particularly its persistent sessions and context handling. Because Harness ran on my computer, its session could outlive a disposable rental:

Day 1: Local Harness session
GPU Server A
Destroy A
Day 2: Same session + Server B

Compaction: remembering the shape, not every token

Long agent sessions accumulate files, terminal output, tests, logs, and messages. Compaction replaces older raw history with a compressed summary or checkpoint while retaining recent detail. It is closer to keeping meeting notes than a full transcript. The tradeoff is important: old facts may survive only in summarized form; the model does not remember every original token verbatim.

Old raw history
Summary / checkpoint
Recent detailed context

Context limits are not output limits

I also encountered output token limits. A context limit means the conversation, tools, and response no longer fit in the total window. An output limit means one generated response becomes too large. Asking the agent to continue sometimes worked; sometimes it repeated part of its earlier answer. Commercial frontier coding agents still felt more polished in some long-horizon situations.

OpenCode

OpenCode is another open-source coding-agent environment that can connect to an OpenAI-compatible provider. I tested its repository tools and context management and found it useful, though I personally preferred DeepSeek Harness for some autonomous workflows. That is a workflow preference, not objective superiority. OpenCode’s configuration changes across versions, so check the documentation installed for your version instead of copying old JSON from a tutorial.

Reproduce the setup

A beginner-friendly path to trying it yourself

Your laptop / desktop
Harness / OpenCode
Ollama /v1
Qwen on rented GPU
  1. Rent an Ubuntu NVIDIA GPU server and connect over SSH.
  2. Clone the setup repository and enter its directory.
  3. Run the 128K Q4 bootstrap unless you have a specific reason to choose BF16.
  4. Verify the API and loaded model.
  5. Connect your harness through a protected endpoint or SSH tunnel.

Clone

git clone https://github.com/khosrotabar/local-ai-coding-box.git
cd local-ai-coding-box

Recommended price/performance

./bootstrap.sh 128 quant

Other supported choices

./bootstrap.sh 128 full
./bootstrap.sh 256 full
./bootstrap.sh 256 quant
./bootstrap.sh auto quant

Check the endpoint and running model

curl http://localhost:11434/v1/models
ollama ps

Start with Q4. BF16 makes sense when you have sufficient VRAM, specifically value the extra precision, and accept the extra cost. Do not begin with fine-tuning. First evaluate the base model, prompting, context size, agent tooling, and your actual work.

Terminology

“Local” here really means self-hosted inference

I controlled the model and inference stack, but I did not always own the physical GPU. Clore and Vast supplied rented hardware. That is different from a workstation under my desk, yet also different from a managed token API.

The fundamental billing unit moved from input tokens + cached tokens + output tokens to primarily GPU time. That shift is powerful for sustained, high-volume work, but it also makes utilization my responsibility. An idle rented GPU still costs money.

What I would do today

The setup I would recommend—and its limits

For most self-hosted coding work, I would begin with Qwen3.8-27B Q4, Ollama, roughly 96K–128K context, and a rented RTX 3090 or RTX 5090 depending on budget. I would attach DeepSeek Harness or OpenCode, test genuine work, and only then consider bigger context, BF16, or fine-tuning.

My favorite setup

Qwen3.8-27B Q4
+ Ollama
+ 128K context
+ rented RTX 5090
+ DeepSeek Harness / OpenCode

In favorable workloads I observed about 120 generated tokens per second on one RTX 5090 rented for roughly $0.26 per hour. At 160 hours, that rate becomes $41.60. The substantial first version of my accounting app cost about $0.80 in GPU rental time. Those are my observed speeds and prices—not guarantees or universal benchmarks.

Open-weight models are now genuinely useful for professional development. Q4 changed the hardware economics, Ollama simplified deployment, and coding harnesses proved nearly as important as the model itself. Context, GPU choice, and automation all mattered. Codex remains in my workflow; Qwen complements it rather than replacing it. Given how quickly open models are improving, this balance may keep changing faster than our tooling habits do.

Sources & tools

References

Hardware throughput, availability, and rental pricing change frequently. Measurements in this article describe my own sessions and should be treated as field observations.

Continue exploring

More engineering field notes are coming.

Frontend security, authorization systems, and practical AI experiments.

All articles →