Skip to main content

Command Palette

Search for a command to run...

What Is an LLM? A Working Model for Engineers

Updated
3 min readView as Markdown

Originally published on DevToolHub.

An LLM is a model that has one job: given the text so far, predict the next token, append it, and repeat. Everything else — chat interfaces, coding assistants, agents — is scaffolding around that one loop.

What is an LLM?

A large language model is a neural network trained to predict the next token in a sequence. Hugging Face's generation docs put it plainly: an LLM "is trained to generate the next word (token) given some initial text (prompt) along with its own generated outputs up to a predefined length or when it reaches an end-of-sequence (EOS) token." There's no separate "understanding" step — just learned statistical likelihood at a scale where it starts looking like reasoning.

How does an LLM generate text?

One token at a time, based on everything before it, using a decoding strategy: greedy search (always the most likely token — deterministic, repetitive) or sampling (weighted by probability, controlled by temperature). This is also why LLMs hallucinate confidently — the model is continuing a pattern, not checking facts.

What is a token, and why does the context window matter?

A token is roughly a word or word-piece. The context window is the hard cap on how many tokens — prompt plus response — the model can hold at once. Ollama's docs disagree with themselves on the default: the FAQ says a flat 4096, the context-length docs say it actually scales with VRAM (4k/32k/256k tiers). Go over it and the model silently drops the oldest tokens. See Ollama hardware requirements for how a bigger context window changes your RAM/VRAM math.

Measuring throughput yourself

ollama run llama3.1 --verbose "Explain quicksort in one paragraph"

This surfaces real prompt eval rate and eval rate numbers in tokens/sec. The API exposes the same as eval_count / eval_duration (nanoseconds) on every call — calculate tokens/sec with eval_count / eval_duration * 10^9. Full request format in the Ollama API guide.

Size, quantization, and hardware

A 70B model needs roughly 8x the memory of a 7B model for better reasoning, at a fraction of the speed. Quantization cuts memory ~4x by lowering weight precision — an 8B Q4_K_M model downloads at ~5GB instead of 16+.

Common mistakes

Defaulting to the biggest model available instead of benchmarking the actual task. Not budgeting context for tool-calling agents — every MCP tool description and response counts against the same window. Assuming greedy decoding is always the safe choice. Skipping the auth model on tool-calling setups — this broke real integrations when MCP's auth spec changed.

Full article with FAQ: devtoolhub.com/what-is-an-llm