speedtest.it
← Blog

Local AI in 8 Milliseconds: A 2B Model Plays Snake on an RTX 5070 Ti

A 2B Decision Model running locally with EuLLM makes decisions in about 8 ms on an RTX 5070 Ti. What it means, and why it's not just about Snake.

Local AI in 8 Milliseconds: A 2B Model Plays Snake on an RTX 5070 Ti

We're used to measuring language models by counting how many tokens they can generate per second.

But there's another interesting scenario.

What if we didn't ask the model to write?

What if it simply had to observe a situation and decide what to do?

A demo built with EuLLM shows a small, Jev-style Decision Model with 2 billion parameters playing Snake entirely locally on an NVIDIA GeForce RTX 5070 Ti.

The number that immediately stands out is the latency: about 8 milliseconds per decision.

But that number only becomes interesting once we understand what's actually being measured.

An LLM playing Snake

The mechanics are simple enough to watch.

On every turn, the model receives information about the game state and has to choose the next direction.

It's not controlling Snake through a traditional algorithm already programmed to follow the optimal path.

The model itself evaluates the situation and picks the action.

Demo: 2B Decision Model plays Snake locally, EuLLM on an RTX 5070 Ti

The response, though, is tiny.

It doesn't have to generate fifty lines of text.

It essentially has to pick from a handful of possibilities.

That radically changes the workload compared to a typical chatbot.

8 ms on a consumer GPU

The machine used for the demo runs an RTX 5070 Ti with 16 GB of VRAM.

Inference runs through EuLLM and doesn't use any external API.

The quantized model is small enough to be kept and run directly on the GPU.

The latency the demo reports is around 8 ms per decision.

That doesn't mean any 2B LLM will always respond in 8 ms.

Latency and throughput depend on the model, the quantization, the input size, the software used, the hardware, and the type of output requested.

Still, the number matters because it shows how different running a model can be when the problem isn't generating long sequences of tokens.

And that 95%?

The demo also shows a second metric: about 95% of the decisions match the best action indicated by a deterministic evaluator.

It's important to read that number correctly.

It doesn't mean the model is "95% intelligent," nor that it has 95% accuracy on any decision-making problem.

There's a separate algorithm that knows the rules of the test and computes which action it considers best.

The model's choices are compared against that reference.

In roughly 95 out of 100 cases, they match.

So it's a metric tied specifically to this demo.

And it's precisely that comparison that makes it more interesting than just a video of an AI playing a game.

Why use a model when an algorithm already exists?

For Snake, a good traditional algorithm is enough and probably is the better solution.

The demo isn't trying to replace it.

It exists to visually show a different category of applications.

Think of a system that receives a technical support ticket written in natural language.

It has to decide whether to route it to networking, security, application support, or a human operator.

Writing out every possible rule by hand quickly becomes unmanageable.

A small, specialized model can interpret the content and simply return the destination.

The same principle applies to AI agents.

An agent might have five tools available and needs to choose which one to use.

Or a RAG system might retrieve a few documents and has to decide whether the information is enough to answer.

These are very different problems from Snake.

But the structure is surprisingly similar:

state → evaluation → decision → action.

Why small models are becoming interesting

For a long time, AI model evolution seemed to follow one direction only: more parameters.

Bigger models have genuinely brought impressive improvements.

But using the biggest available model for every problem isn't necessarily efficient.

When the domain is narrow and the possible actions are limited, a specialized model can be enough.

And a small model comes with very concrete advantages.

It needs less memory.

It can run on relatively common hardware.

It reduces inference time.

And, above all, it can run locally.

No trip to the cloud

In the EuLLM test, the data doesn't leave the GPU, cross the internet, reach a provider's datacenter, and come back with an answer.

The entire process happens on the machine.

That difference matters especially when an application has to run many decisions in a row.

Even a very fast API has to deal with network latency.

Ten or twenty calls within the same workflow can turn a few hundredths of a second into several seconds the user can actually notice.

Local inference removes that part of the path.

It doesn't eliminate the computational cost, of course — the hardware still has to run the model.

But it moves the problem from the network to the local machine.

Not just datacenter GPUs

One of the most interesting aspects of the demo is precisely the hardware.

The RTX 5070 Ti isn't a professional datacenter GPU.

It's a consumer card.

Sixteen gigabytes of VRAM still puts real limits on how large a model can be to fit entirely on the GPU, but a few-billion-parameter quantized model fits comfortably into the kind of workload that becomes interesting on this class of hardware.

It's a signal of how the local AI market is shifting.

The comparison is no longer just workstations with GPUs worth thousands of euros against cloud services.

There's now a middle tier made of small servers, workstations, and even some mini-PCs with large amounts of shared memory.

Decision Model: a term we'll hear more often?

The term doesn't describe a fundamentally new architecture like Transformer or Mixture of Experts.

It describes the role assigned to the model instead.

A Decision Model receives a state or context and mainly produces a choice.

Language generation can be minimal or even secondary.

The idea is becoming particularly interesting alongside the growth of AI agents, because an agent spends much of its time precisely deciding what to do next.

If that decision can be handed to a small local model instead of constantly querying a huge cloud model, costs, latency, and the architecture of the whole system change.

Snake is just the easiest way to see it

A demo has to make what's happening immediately clear.

Snake does that very well.

Does the model decide badly? We see it right away.

Does it decide correctly? The snake keeps going.

But the interesting part comes when we swap the game's four directions for actions like:

approve, reject, search again, call a tool, hand off to a human operator.

At that point we're no longer talking about video games.

We're talking about how part of the AI software of the next few years might be built.


The EuLLM demo and code are available on the project's public repository: github.com/eullm/eullm.


← All articles