The Usable Size of a Context Window

By Francis Nguyen, Chief Executive Officer··Research
A light monochrome chart of nested squares titled "What fits, and what is usable," and subtitled "On o200k, a one-million-token window holds 4.3 to 5.0 MB of English. Harder tasks use a corner of it," with area proportional to bytes. The large outer square is what one million tokens of English holds, 4.3 to 5.0 MB, measured by Ardenus on OpenAI’s o200k tokenizer. Inside it, sharing one corner, three smaller and darker squares mark how much a model can reliably use by task, an Ardenus synthesis of published benchmarks drawn at the top of each range: 0.55 to 0.65 MB to find and order several similar items, for the best models, at about 128,000 tokens; 34 to 160 KB to count or aggregate across records, at 8,000 to 32,000 tokens; and up to 69 KB to infer across facts that share few words, at 16,000 tokens or less. A right rail adds that one gigabyte of English is 198 to 232 million tokens, 198 to 232 such windows end to end, and that on Claude’s tokenizer since Opus 4.7 one million tokens holds an estimated 2.6 to 3.0 MB of English.

Introduction

How much data can a large language model work with before it loses track of the evidence or starts inventing it? Vendors answer in tokens, the chunks of text a model reads and is billed by, and as of October 2026 the frontier standard is a window of about one million of them. In the units an operator plans with, that window holds a few megabytes of text, not a gigabyte, and how much of it a model can reliably use depends on the task more than on the window: all of it to find one fact by its wording, in the vendors’ own tests; on independent tests, hundreds of kilobytes to find and order several similar items on the strongest current models and tens to low hundreds of kilobytes to count across records; and, for the 2024 and 2025 models tested, tens of kilobytes at most to connect facts that share few words. We measured what fits ourselves, on public text, and assembled what is usable from the published evidence. A disclosure up front, because it matters for how you read this: Ardenus sells a data layer that retrieves and computes over an operator’s records instead of pasting them into a model’s window, so a finding that usable context is small is convenient for us. Weigh it knowing that.

  • A million tokens is megabytes, not gigabytes. On the tokenizer behind OpenAI’s GPT-5 models it holds 4.3 to 5.0 MB of English prose, 3.7 to 4.7 MB of source code, and 2.0 to 3.3 MB of structured records. One gigabyte of English is 198 to 232 million tokens.
  • Finding one fact by its wording works across the whole window in the vendors’ own tests. Finding several similar items, counting across records, or connecting facts that share few words does not: on independent tests the reliable zone is hundreds of kilobytes at most, and for inference it is tens of kilobytes at most, for the 2024 and 2025 models tested.
  • Fabrication generally rises with length where it has been measured. In the largest such study we found with the answer absent by construction, which covers open models only, the lowest rate was 1.19 percent on 115 KB of documents, and every one of the 11 models run on 728 KB fabricated more than 10 percent of the time. The only length series we found that could include closed frontier models does not break its results out by model, so we found no per-model curve for any closed frontier model.
  • The usable size is a product, not a window. It is the length your model can reliably handle for your task, multiplied by the bytes each token of your data carries, and for questions that compare, count, or connect records it is far smaller than the advertised window.

What a window holds

A token is not a fixed amount of data. Vendors offer a rule of thumb, roughly four characters of English per token in OpenAI’s documentation and Google’s, but the real ratio depends on the tokenizer and, far more, on what the text is. So we measured it. In two passes we ran public text through fifteen published tokenizer files, covering ten distinct tokenizers, and counted how many bytes of plain text each token carries: two Project Gutenberg novels, Wikipedia articles, open-source code in five languages, server logs, public city records, taxi-trip records, economic time series, and the FLORES-200 parallel sentences used to compare languages. An independent re-implementation, written to audit the pipeline, matches it to the token on the same book, Jane Austen’s Pride and Prejudice: 159,225 tokens for 734,437 bytes. The pipeline also reproduces the published language premiums of Petrov and colleagues for all nine languages we checked, to two decimal places. On o200k, the tokenizer that OpenAI’s tiktoken library maps to GPT-4o through the GPT-5 family, one million tokens holds 4.3 to 5.0 MB of English prose, about 760,000 words. OpenAI does not name a tokenizer for its GPT-6 models, so for them o200k is a stand-in. The vendors’ rule of thumb is, if anything, conservative for English.

A horizontal bar chart titled "What one million tokens holds": megabytes of each kind of content per one million tokens of OpenAI’s o200k tokenizer, each bar solid to the low end of the range measured across samples and lighter to the high end. English prose 4.3 to 5.0 MB; source code 3.7 to 4.7 MB; server logs 2.0 to 3.3 MB; records as minified JSON 2.9 to 3.3 MB; records as pretty-printed JSON 2.6 to 2.9 MB; records as CSV 2.0 to 2.5 MB; records as a Markdown table 2.2 to 2.3 MB; a numeric time series as CSV 1.5 to 1.6 MB; and a binary file pasted as base64 1.1 MB of the underlying file. An Ardenus measurement on public text, in two passes.
Megabytes of each kind of content in one million tokens of OpenAI’s o200k tokenizer; for base64, megabytes of the original binary file. Each bar runs solid to the low end of the range we measured across samples, with each server log counted as its own sample, and lighter to the high end. Ardenus measurement on public text, in two passes; 1 MB is 1,000,000 bytes.

Almost everything else holds less. The same million tokens carries 3.7 to 4.7 MB of source code, 2.0 to 3.3 MB of server logs, 2.0 to 3.3 MB of structured records depending on the format, 1.5 to 1.6 MB of a numeric time series, and only 1.1 MB of a binary file pasted in as base64 text. Among the tokenizers we ran ourselves, the choice barely matters for clean English prose: they differ by 2 to 5 percent on any one cleaned sample, though by 13 percent on one novel left exactly as distributed. It matters a great deal for numbers. Tokenizers that split every digit into its own token, such as Qwen, Gemma, and Mistral’s Tekken, spend 25 to 26 percent more tokens than o200k on a CSV of taxi trips with short numbers, mostly the parts of dates and times, 47 to 58 percent more on a numeric time series, and 67 to 73 percent more on a CSV of inspection records full of long identifiers and coordinates.

The tokenizer also moves the answer for Claude. Anthropic’s newer models count text differently: on the same English text, the tokenizer Claude has used since Opus 4.7 spends about 1.66 tokens for every o200k token. That figure is our own re-analysis of the token-efficiency ratios, built from API-reported token counts, that the independent Context Arena leaderboard publishes for Google DeepMind’s released benchmark text, so it is an indirect estimate, but it implies that a one-million-token window on Claude Opus 4.7 and later models holds about 2.6 to 3.0 MB of English, not 4 to 5. Anthropic’s own documentation puts it a little lower, at “roughly 555k words or 2.5M Unicode characters on the current tokenizer.” Language moves it further. The same content in Hindi costs 1.56 times the English tokens on o200k and 4.78 times on the older tokenizer behind GPT-4, or 4.79 on the larger sentence set Petrov and colleagues used; because Hindi script takes three bytes per character, a million o200k tokens is 8.08 MB of Hindi yet carries only the content of about 641,000 English tokens.

The windows on sale have converged on roughly the same size. As of October 2026, Anthropic’s current models offer one million tokens, with up to 128,000 tokens of output counted inside the window; OpenAI’s GPT-6 models list 1,050,000 tokens with up to 128,000 of output; Google’s Gemini 3.1 Pro Preview and Gemini 3.8 Flash take 1,048,576 input tokens. The two-million-token API windows are gone: Google’s Gemini 1.5 Pro no longer appears in its current documentation, and xAI retired its 2M-window API models on May 15, 2026. The outlier is Meta’s open-weight Llama 4 Scout, advertised at ten million tokens, roughly 42 to 51 MB of English, though Meta trained it at 256,000. None of these reach a gigabyte. Even Magic’s 100-million-token model, announced in 2024 with no release posted on Magic’s blog since, would hold about half of one if its tokens match o200k’s; Magic has not disclosed its tokenizer.

How many records fit in a context window?

Operational data is mostly records, not prose, and the format a record is written in decides how many fit. We serialized 5,000 public New York City rodent-inspection records, 23 fields each and the same public dataset behind our rodent-inspection research, in six formats with identical fields, plus the city’s own CSV export; the other pass serialized 6,433 New York City taxi trips, 14 fields each, from the public seaborn-data taxis table, in six formats, likewise with identical fields. Per field value, pretty-printed JSON costs about 10.0 to 10.9 tokens, minified JSON 6.9 to 7.6, a Markdown table 4.6 to 5.3, and CSV 3.9 to 4.8. A million tokens therefore holds about 4,400 to 6,600 records as pretty-printed JSON, or 11,300 to 14,800 as CSV, and the same data costs 2.26 to 2.58 times as many tokens as pretty-printed JSON as it does as CSV. The keys repeat on every row, and every repetition is paid for.

File size is not context capacity either. The text we extracted from three real PDFs comes to only 2 to 19 percent of their file size, at 676 to 1,319 o200k tokens of text per page. Vendors bill most non-text input by geometry and time, not bytes: Anthropic’s PDF support says each page “typically uses 1,500–3,000 tokens per page depending on content density” for its text, plus image tokens because each page is also converted into an image, and a Gemini model with a one-million-token window can take about three hours of video at low resolution or one hour at high. Spreadsheet uploads can be reduced before the model sees them: OpenAI’s file inputs parse at most the first 1,000 rows of each sheet and add a model-generated summary. And request limits are payload limits, not capacity: Anthropic’s file documentation gives a 500 MB plain-text file as an example of something larger than the context window.

Fitting is not using

Whether a model can use what fits depends on what it is asked to do, and the published evidence sorts tasks into a rough order. The easiest is finding one literal fact. In OpenAI’s internal evaluation, GPT-4.1 found a single planted fact at every position up to one million tokens, and Google reported similar results for Gemini 1.5 Pro in 2024. Those are vendor-run, single-needle tests: they show a real capability, and why a capability score is not the same thing as value in an operation is the subject of our essay on the limits of benchmarks as value signals. For pure lookup, the whole window works.

Finding several similar items and keeping them in order is harder. OpenAI’s MRCR benchmark, short for multi-round co-reference resolution, hides eight identical requests in a long synthetic conversation and asks the model to return a specific one by its position among them. We measured the released benchmark itself: its longest length bin, labelled “1M,” averages 3.88 MB of text, between 2.6 and 5.2 MB per test. As of October 2026, vendor-published scores in that bin split widely. OpenAI reports GPT-5.5 at 74.0 percent at its highest reasoning setting, and Anthropic reports Claude Opus 4.6 at 76.0 percent at maximum effort and 78.3 percent with 64K extended thinking. Claude Opus 4.7 scores 32.2 percent at maximum effort, read from a chart in its system card, but that is not a clean comparison: by our indirect estimate of Claude’s newer tokenizer, the average prompt in this bin runs past Opus 4.7’s one-million-token window, and the card does not say how such prompts were handled. In the same bin, Context Arena’s archived board for OpenAI’s test puts Gemini 3.1 Pro at 25.9 percent, an independent run. The highest scores are both self-reported, and no one has replicated them: 96.3 percent for OpenAI’s GPT-6 Astra, its best result across effort settings, and 98.1 percent for Meta’s Muse Spark 1.3 at maximum effort, read from a scorecard image. Context Arena’s current board runs Google DeepMind’s separate version of the test, so its numbers cannot be set beside OpenAI’s. On it, the strongest current runs score about 86 to 92 at 128,000 tokens and 30 to 78 at 512,000, and no OpenAI or Anthropic model has a result at one million. On independent evidence, the reliable zone for this task is about 128,000 tokens, or 0.55 to 0.65 MB of English, for the strongest current models, and well short of that for many others.

Inference collapses earlier, at least for the 2024 and 2025 models tested so far. The NoLiMa benchmark plants facts that share few words with the question, so the model must connect them rather than match them, and measures the length at which a model keeps 85 percent of its short-context score. The published effective lengths are 16,000 tokens for GPT-4.1, about 69 KB of the benchmark’s book text by our measurement; 8,000 for GPT-4o, about 34 KB; and 1,000 for Llama 4 Scout, a few kilobytes, despite its ten-million-token window. Reasoning models degrade too: on the benchmark’s harder set, o3 fell from 100 to 58.5 at 32,000 tokens. Why models use long inputs unevenly, and fall short of their advertised windows, is covered in the data foundation beneath enterprise AI.

Counting and aggregating degrade steadily. On the synthetic split of Oolong, a benchmark of counting and classifying many short records, most models score below 50 at 128,000 tokens, and GPT-5 falls from 85.56 at 8,000 tokens to 46.36 at 128,000. On TQA-Bench, a multiple-choice test over real database tables where chance is 25 percent, GPT-4o’s lookups hold steady while its sums fall from 76.00 to 32.63 as the tables grow from the benchmark’s setting labelled 8K to the one labelled 64K, which is about 41,700 tokens in practice. Older evidence on a different task points the same way at even smaller sizes: for five models of late 2023, average accuracy on true-or-false questions that depend on two facts buried in padding fell from 0.92 to 0.68 by about 3,000 tokens, in a study by Levy, Jacoby, and Goldberg.

The operating envelope

Multiply the effective lengths by the bytes per token and the usable size of a window comes out in units an operator can plan with. The table below is our synthesis, not a measurement: it combines the published thresholds above with our byte counts, it draws each zone at the generous end of its range, and individual models sit well above or below it. The inference and fabrication rows use each benchmark’s own text as we measured it, and records are counted at the 67 to 89 o200k tokens per CSV record we measured.

TaskWhere the published evidence puts the limitEnglish proseNumber of CSV records
Find one literal factThe full window, in vendor tests to one million tokens4.3 to 5.0 MBAbout 11,300 to 14,800
Find and order several similar itemsAbout 128,000 tokens for the strongest current models on independent tests0.55 to 0.65 MBAbout 1,400 to 1,900
Count, sum, or aggregate across recordsDegradation visible from 8,000 to 32,000 tokens34 to 160 KBAbout 90 to 475
Infer across facts that share few words16,000 tokens or less for 2024 and 2025 modelsAbout 69 KB or lessNot measured
Decline to answer when the evidence is absentFabrication above 10 percent for every open model tested at 728 KB115 to 728 KB testedNot measured

Read the table as orders of magnitude. Lookup lives in megabytes; comparison in hundreds of kilobytes; counting in tens to low hundreds of kilobytes, which is a few hundred records, not a few thousand; and inference in tens of kilobytes or less. Many newer models push the boundaries outward, though not reliably, and a team that evaluates its own model on its own task should trust that result over this table.

Losing context or inventing

Two different failures hide inside “the model got it wrong.” Losing context means the evidence was present and went unused. Inventing means the model asserted something its input does not support. The published evidence on the second is much thinner than on the first. The largest length-stratified fabrication study we found with the answer absent by construction, RIKER2, a 2026 preprint from Kamiwaza AI, tests 35 open-weight models on synthetic document sets with probe questions about entities and fields the documents do not contain, where the correct response is to say so. The lowest fabrication rate any model reached was 1.19 percent at the study’s nominal 32,000-token size and 3.19 percent at 128,000, both by GLM 4.5, which was not run at the largest size; at 200,000, all 11 models tested fabricated more than 10 percent of the time, the best at 10.25 percent. We measured the released document sets: 115 KB, 411 KB, and 728 KB of text. The rise is not uniform, and several models did better at the middle size than the smallest.

Summaries show the same direction. In Vectara’s 2025 leaderboard data, read from a chart that does not break out models and judged by its own hallucination model, the hallucination rate rises from 6.5 percent for source articles under about 1,550 words to 35.3 percent for articles of 13,594 to 15,099 words, roughly 78 to 100 KB at the bytes per word we measured for English prose, though it plateaus near 20 to 22 percent before that last jump, and length is not separated from how complex the articles are. How models fail can differ by family, but it depends on the test set. In Chroma’s needle tests with distractors the needle was always present, so an abstention there, declining to answer, is lost context, not a correct refusal. On one test set, read from its charts, 73.7 percent of Claude models’ failures were abstentions, against 4.9 percent for GPT models, whose failures were overwhelmingly answers taken from a distractor. A second set split the same way, a third far more narrowly, and on a fourth every family’s failures were abstentions. What we did not find is a fabrication curve for any closed frontier model on its own; Vectara’s series does not separate models. None of the vendor documents we checked reports such a curve. Anthropic’s system cards for Claude Opus 4.8 through Haiku 5.5 report long context only through agentic coding, plus graph traversal in some of them; Google’s Gemini 3.1 Pro evaluation note only through retrieval, on its own version of the multi-item test; and OpenAI’s GPT-6 Astra system card evaluates hallucination only on ChatGPT conversations that users had flagged for factual errors, with no breakdown by length. The broader case for tying a model’s answers to evidence is in grounding and the reliability of AI.

Where the gigabytes are

If no window holds a gigabyte of data, the gigabytes still show up, on the hardware side. A model keeps a working cache, which engineers call the key-value cache, for every token in its window, and that memory grows with length. From Llama 3.1 70B’s published configuration and the standard cache formula, counted over its 8 key-value heads rather than its 64 query heads, each token costs 327,680 bytes at 16-bit precision: about 43 GB of GPU memory for one sequence that fills the model’s real window of 131,072 tokens, holding 0.56 to 0.65 MB of English. Stretched to a hypothetical million tokens, the same design would need about 328 GB to hold 4.3 to 5.0 MB of text, some 66,000 to 77,000 times the bytes it represents. Older designs were worse: KVQuant reports 512 GiB of cache at 1,048,576 tokens, about 524 GB per million, for the original LLaMA-7B, which stored full keys and values for every attention head. Time grows too: filling a million-token prompt took 30 minutes for an 8-billion-parameter model on one A100 in the MInference study, a cost the study’s own sparse-attention method cut by up to ten times. And the price per token rises past a threshold at most vendors. As of October 2026, OpenAI charges twice the input rate and 1.5 times the output rate for the whole request once a prompt exceeds 272,000 tokens on its million-token models, Google does the same above 200,000 on its Pro models, xAI bills every token at about twice the rate once a prompt reaches 200,000, and Anthropic bills its full window at standard rates except for Haiku 5.5, which charges five times as much once a prompt passes 100,000 tokens.

Where the argument bends

The case for small usable context has real limits, and the strongest evidence against it deserves its full weight. Single-fact lookup genuinely works across a million tokens in the vendors’ own tests. The two self-reported results of 96 to 98 percent on OpenAI’s multi-item test, if they replicate, move the comparison boundary out by close to an order of magnitude for those models. Some models hold up where most do not: Gemini 3.1 Pro Preview stayed flat at about 96 to 97 percent through TQA-Bench’s largest setting, Gemini 3 Pro still scored 62.42 on Oolong’s synthetic split at 128,000 tokens, down from 90.94 at 8,000 but above the 50 that most models fell below, and on the Fiction.LiveBench story-comprehension test, read from its April 2026 results table, GPT-5.2 scored 96.9 and Gemini 3 Flash Preview 100 at 192,000 tokens. Length is not always the culprit either: a single-author 2026 preprint found that cutting inputs to a quarter while keeping every fact the answer needs left Claude Opus 4.7 and GPT-5.5 with no significant change on two long-context tests of up to 32,000 and 200,000 tokens, while two smaller Claude models, Haiku 4.5 and Sonnet 4.6, improved significantly on the shorter one, so for the strongest models what fills the window can matter more than how much of it there is. And newer is not reliably better: in the 128K-to-256K bin of OpenAI’s test, Claude Opus 4.7 scored 59.2 against Opus 4.6’s 93.0, both at maximum effort as Anthropic reports them, with the 59.2 read from a chart in the Opus 4.7 system card. The drop does not carry over to GraphWalks, a graph-traversal test on which Anthropic’s Opus 4.8 system card shows Opus 4.7 ahead of Opus 4.6 on three of four scores.

Our own numbers have limits as well. We measured Claude’s and Gemini’s tokenizers only indirectly, Claude’s through the token-efficiency ratios Context Arena publishes from API-reported counts and Gemini’s through the token lengths Google DeepMind released with its benchmark text, and only on English. The newest models, including Claude Opus 5.5, the GPT-6 family, and Gemini 3.8 Flash, have no independent long-context results that we could find. And the operating envelope is a synthesis across benchmarks with different thresholds, not a single measurement.

What this means for the operator

The practical test is simple. Count your rows, check what format they will be sent in, and multiply by the tokens each record costs in that format: on o200k, we measured 67 to 89 per record as CSV for records of 14 to 23 fields, and pretty-printed JSON costs 2.26 to 2.58 times as many. Then match each question you want answered to the right row of the envelope. “Find the inspection for this work order” is lookup, and a window of records can serve it. “Find this customer’s last inspection” sounds like lookup but is comparison, because the model has to find every inspection on the account and work out which came last; so is “which of these eight similar complaints at this address came first.” Both hold for about 1,400 to 1,900 records as CSV at best, and only on the strongest current models. “How many inspections failed last quarter, by branch” is counting, and past a few hundred records it belongs in a database query, not a prompt. Once you know the size of the usable window, the argument for spending less of it is in the diminishing returns of context, and predictions from historical records belong to models built for records, as the regimes of prediction sets out.

The disclosure from the top belongs here in full, because the argument points at what Ardenus sells: a governed data and intelligence layer that unifies the records from the systems an operator already runs and hands a model only the slice a question needs, so the finding that usable context is small is seller-convenient, and a reader should weigh it knowing that. It is fair to say plainly what the other side has: lookup across a million tokens works in the vendors’ own tests, the best models are improving on harder tasks, and a team that pastes a small, well-chosen slice into a long window is doing nothing wrong. No result, saving, or metric for Ardenus’s products or clients is claimed here, and no client data appears anywhere in this essay: every byte count comes from public text, and every token count from a public tokenizer except the Claude estimate, which rests on Context Arena’s published token-efficiency ratios; the methodology note below describes the inputs. You can read more of our research on the Ardenus articles hub, or see the platform on the technology page.

Sources and methodology

The “fits” numbers are our own measurement. We counted the UTF-8 bytes that each token of public text carries, in two passes. The first ran fifteen published tokenizer files (ten distinct tokenizers, including OpenAI’s o200k_base and cl100k_base and the Llama, Qwen, DeepSeek, Gemma, and Mistral tokenizers) over pinned public inputs: two Project Gutenberg novels, source code from CPython, React, tokio, gin, and zod, sixteen Loghub server logs, 6,433 New York City taxi-trip records in six formats, the FLORES-200 parallel sentences, and a PDF pasted as base64. The second ran eight of those tokenizers over a partly shared corpus adding Wikipedia text, a raw web page, 5,000 New York City rodent-inspection records in six formats plus the city’s CSV export, Federal Reserve economic time series, and three public PDFs measured for text against file size. The passes share some inputs and code, so they are not independent of each other; an independent re-implementation of the count matches both on the same book to the token, and the fifteen-tokenizer pass reproduces Petrov and colleagues’ published premiums for nine languages to two decimal places. We also measured the released text of OpenAI’s MRCR benchmark, NoLiMa’s haystack books, Google DeepMind’s GDM-MRCR benchmark and the RIKER2 document sets, so each benchmark’s lengths convert to bytes in its own text. Claude’s tokenizer was measured only indirectly, from the token-efficiency ratios Context Arena builds from API-reported token counts. The “usable” numbers are published thresholds, each re-verified against its primary source in an adversarial fact-check of our research brief (119 figures checked across the brief: 91 confirmed, 28 corrected, none rejected), with vendor results labelled as vendor-reported. The operating envelope is our synthesis of the two. No client data was used.

  1. OpenAI MRCR (OpenAI, Hugging Face dataset) - eight-needle multi-round co-reference resolution; its 512K to 1M bin averages 3.88 MB of text by our measurement.
  2. Introducing GPT-5.5 and GPT-6 Astra (OpenAI, 2026) - vendor-reported MRCR scores, the latter the maximum at any effort.
  3. Claude Opus 4.6 System Card, Claude Opus 4.7 System Card, and Claude Opus 4.8 System Card (Anthropic, 2026) - vendor-reported MRCR and GraphWalks scores.
  4. Introducing Muse Spark 1.3 (Meta, 2026) - self-reported MRCR scores on OpenAI’s data, read from a scorecard image.
  5. Context Arena (independent leaderboard) - GDM-MRCR v2 runs by bin and token-efficiency ratios built from API-reported token counts; its archived OpenAI-MRCR board is the source of the independent Gemini 3.1 Pro run, 25.9 percent in the 1M bin, which the Claude Opus 4.7 system card also charts.
  6. Introducing GPT-4.1 in the API (OpenAI, 2025) - single-needle retrieval at every position to one million tokens, an internal evaluation.
  7. NoLiMa: Long-Context Evaluation Beyond Literal Matching (Modarressi et al., ICML 2025) and its repository updates - effective lengths when facts share few words with the question.
  8. Oolong: Evaluating Long Context Reasoning and Aggregation Capabilities (Bertsch et al., COLM 2026) - counting and aggregation by length.
  9. TQA-Bench: Evaluating LLMs for Multi-Table Question Answering (Qiu et al., accepted to IEEE Transactions on Big Data) - lookup versus counting over tables.
  10. Same Task, More Tokens (Levy, Jacoby & Goldberg, ACL 2024) - reasoning accuracy as padding grows.
  11. How Much Do LLMs Hallucinate in Document Q&A Scenarios? (Roig, Kamiwaza AI, 2026 preprint) - fabrication by input length for open-weight models.
  12. Introducing the Next Generation of Vectara’s Hallucination Leaderboard (Vectara, 2025) - summarization hallucination by source length, not broken out by model.
  13. Context Rot (Hong, Troynikov & Huber, Chroma, 2025) - failure style by model family and test set.
  14. Checked for fabrication by input length, none reporting it: System Card: Claude Sonnet 5.5 (Anthropic, 2026, one of the cards from Claude Opus 4.8 through Haiku 5.5 that we checked), GPT-6 Astra System Card: Hallucinations (OpenAI, 2026), and Model Evaluation: Gemini 3.1 Pro (Google DeepMind, 2026).
  15. Fiction.LiveBench, April 2026 (Fiction.live) - story comprehension by length, values read from a results image; and Distractor-Aware Truncation (Arjmandi, 2026 preprint) - cutting inputs while keeping every answer-bearing fact.
  16. Language Model Tokenizers Introduce Unfairness Between Languages (Petrov et al., NeurIPS 2023) - the token premium of the same content across languages.
  17. KVQuant (Hooper et al., NeurIPS 2024) and MInference 1.0 (Jiang et al., NeurIPS 2024) - key-value cache memory and prefill time at long context.
  18. Vendor documentation, as of October 2026: Anthropic context windows, Anthropic models overview, Anthropic pricing, Anthropic PDF support, Anthropic Files API, OpenAI key concepts, OpenAI file inputs, OpenAI GPT-6 Astra model page, OpenAI tiktoken, Gemini token counting, Gemini video understanding, Google Gemini 3.1 Pro, Gemini pricing, xAI models and pricing, xAI model retirement, Meta Llama 4 announcement, Magic, 100M token context windows, and NVIDIA on the key-value cache formula.