The detail
Research papers tokenize heavier than their word count suggests: citation strings, author names, DOIs, statistical notation, and table contents are all token-dense. A "6,000-word" paper extracted from PDF routinely lands between 10,000 and 13,000 tokens.
That still fits comfortably everywhere: about 10% of a 128K window, 5% of 200K, and ~1% of a 1M-token window — which is why literature-review workflows now load dozens of papers into a single long-context request.
PDF extraction quality drives real token counts more than paper length. Two-column layouts, equations, and figure captions can inflate or garble extraction; cleaning extracted text before prompting both cuts tokens and improves answer quality.