It is late, and a question arrives the way they usually do, sideways, from something half-read: what is the difference between a virus and a viroid? I don’t open a browser with its eleven tabs, or one of the large hosted models. I type the question into a window on my own machine, and a few seconds later an answer appears. A viroid is a bare loop of RNA, no protein coat, no genes for making proteins, a parasite reduced almost to a single gesture. The answer is not brilliant. It is adequate, it is immediate, and it never left the room.
The machine is a Mac with an M4 chip and 24 gigabytes of memory, modest by the standards of anyone who trains models and perfectly ordinary by the standards of anyone who doesn’t. It does not write my code. It does not do my research. It cannot compete with the frontier systems on a hard proof or a long synthesis, and I don’t ask it to. I use it the way earlier generations used the dictionary on the shelf and the encyclopedia in the hallway: as the first stop for a question too small to justify a trip.
I. A Dictionary, Not an Oracle
“Small” is a relation, not a verdict. A model that would embarrass itself on a software-engineering benchmark can still explain why the sky is darker blue overhead than at the horizon, what satisficing means, or how a ribosome reads a codon. Herbert Simon coined that word for the decision that is good enough rather than optimal, and he did not mean it as a confession of laziness. Under limited time and attention, he argued, satisficing is what bounded rationality looks like when it works.
What makes the arrangement possible is mostly arithmetic. A model with eight to fourteen billion parameters, quantized from sixteen bits per weight down to about four, fits in five to nine gigabytes, small enough to share memory with everything else a laptop is doing. Projects like llama.cpp and Apple’s MLX turned that compression into something a non-specialist can run in an afternoon. The parameter count is not a full measure of quality, and the quantization costs something, but for definitions, explanations, and first orientations the loss is often smaller than the gain in proximity.
Reference works have never been infallible, and nobody expected them to be. Samuel Johnson’s Dictionary of the English Language (1755) defined pastern as “the knee of a horse.” It is not. When a lady asked him how he could have made such a mistake, Boswell records the answer: “Ignorance, madam, pure ignorance.” The book remained indispensable for a century. Its readers knew what kind of object they were holding.
II. Who Else Is in the Room
The difference between a local model and a hosted one is not intelligence. It is the number of parties to the conversation.
A question sent to a frontier service becomes a record held by a company, governed by its retention policy, stored on its infrastructure, and reachable by whatever law reaches the company. In the United States that reach is wide. The third-party doctrine, built in United States v. Miller (1976) and Smith v. Maryland (1979), holds that information voluntarily handed to a third party carries a reduced expectation of privacy. Carpenter v. United States (2018) narrowed it for cell-phone location data, but did not abolish it. The CLOUD Act of the same year made clear that American providers must produce data in their custody wherever in the world the servers sit. And in 2025, a federal judge in The New York Times v. Microsoft and OpenAI ordered OpenAI to preserve user conversations it would otherwise have deleted, including ones users had deleted themselves.
None of this means that anyone is reading my question about viroids. Most prompts are never seen by a person, and the companies involved publish transparency reports and fight some demands in court. The point is narrower. Every question sent elsewhere asks me to trust a chain of policies, employees, contractors, and jurisdictions I can’t inspect. For a contract, a codebase, or a hard research problem, that trust may be worth extending. For the meaning of a word, it is a strange price to pay by default.
III. The Reference Desk Never Asked Why
Libraries understood this long before software did. The reference interview teaches librarians to clarify what a patron needs without demanding why they need it. In 2005, four Connecticut librarians from a consortium called Library Connection received a National Security Letter asking for patron records and, under a gag order, went to court instead of complying quietly; the case became Doe v. Gonzales. The profession’s instinct was that a question asked at the desk belongs to the person who asked it.
There is evidence that the instinct matters. After Edward Snowden’s 2013 disclosures about programs like PRISM, the legal scholar Jonathon Penney measured traffic to Wikipedia articles on topics like terrorism and security and found a sharp and sustained drop. Nobody had been arrested for reading about Al-Qaeda. People simply became less willing to look. That is what a chilling effect is: not punishment, but the quiet withdrawal of curiosity in anticipation of being seen.
Some questions are banal, some are embarrassing, some touch on health or politics, and some are nobody’s business for no particular reason. A local model lowers the cost of asking all of them. That does not make the questions more important. It makes them easier to ask.
IV. What “Local” Does Not Mean
The word “local” can become its own superstition. The weights sit on my disk and the arithmetic happens on my chip, but the application wrapped around them may still send telemetry, keep a chat history in a plain file, sync that file to a cloud backup, or call a web-search tool when I forget it is switched on. The operating system indexes what it finds. A browser extension can read the window. The NIST profile on generative-AI risk lists exactly these leaks, prompts, logs, and integrations, as places where data escapes systems that look contained.
So the honest claim is modest. Local inference removes one large dependency without removing all of them. Helen Nissenbaum’s idea of contextual integrity gives the right frame: privacy is not secrecy but the expectation that information flows according to the norms of the context it came from. A question asked at my desk has the norms of a desk. The engineering principle that follows is old and unglamorous, data minimization: do not transmit what does not need to be transmitted.
V. The Year on the Spine
The small model has the faults of a reference work, plus a few of its own. Its knowledge stops at a date, the way a printed encyclopedia has a year on its spine. It invents with the same calm voice it uses to report, and it is likelier to invent than a larger model. Ask it about a minor eighteenth-century botanist and it may hand you a birth year from nowhere.
Borges built a story on that risk. In “Tlön, Uqbar, Orbis Tertius”, the narrator’s friend finds an article on a country called Uqbar in one copy of an encyclopedia and in no other copy of the same edition. The article is fluent, specific, cross-referenced, and describes nothing. A language model is a press that can print that volume on demand. Its answer is a map of what people have written, not the territory they were writing about, and, as in Borges’s Library of Babel, the faithful entry and the false one sit on the same shelf in the same binding.
That is why the role stays bounded. The model is good for the first explanation, for naming a question, for telling two nearby ideas apart, and for knowing what to read next. When something matters, a dosage, a date in an argument, a claim I intend to repeat, the answer is a pointer, not a source. I follow it to a book, a paper, or a better tool. The pocket books in my father’s suitcase worked the same way: their value was that they ended, and pointed past their own edge. They had another virtue I only recognize now. Nobody knew what I was reading.
VI. Close to Home
The question I have stopped asking is “which model is smartest?” The useful one is “what is the least powerful tool that serves this question well?” Code, specialized research, and difficult synthesis may justify the frontier and its tradeoffs. A word, a phenomenon, or a half-remembered idea from a podcast usually does not.
What a small model on the desk changes is not the size of my knowledge but the threshold for seeking it. A question that would have dissolved because opening a tab felt like too much effort now gets an answer, imperfect and immediate, while it is still alive. Most of those answers will never matter. A few become the first page of something longer. None of them had to be announced to anyone.
Further reading
- Herbert Simon — satisficing and bounded rationality
- Georgi Gerganov et al. —
llama.cpp - Apple Machine Learning Research — MLX
- Samuel Johnson — A Dictionary of the English Language
- The third-party doctrine, Smith v. Maryland, and Carpenter v. United States
- The CLOUD Act
- The New York Times v. Microsoft and OpenAI
- The reference interview, National Security Letters, and Doe v. Gonzales
- PRISM
- Jonathon W. Penney — “Chilling Effects: Online Surveillance and Wikipedia Use”
- NIST — AI 600-1: Generative Artificial Intelligence Profile
- Helen Nissenbaum — contextual integrity
- Data minimization
- Jorge Luis Borges — “Tlön, Uqbar, Orbis Tertius”
