Long inputs sag in the middle and the fix is structural

Capacity and attention are not the same thing
These tools will accept a great deal of material, and the amount they accept has grown considerably. What has not kept pace is uniform use of it. Material near the beginning and the end of a long input tends to have more influence than material in the middle, and instructions buried inside a large document may be followed loosely or missed.
This shows up in ordinary work as a specific, recognisable pattern. A summary of a long report covers the opening and the conclusion well and thins out across the centre. An instruction given before forty pages of source material is applied at the start of the answer and drifts by the end.
The behaviour varies between tools and versions and is not something to rely on precisely. But the general direction is consistent enough that designing around it is cheaper than testing whether it applies to you today.
Select rather than supply
The first structural response is to stop treating a large capacity as a reason to paste everything. Twelve relevant pages produce better output than four hundred pages containing those twelve, and finding the relevant pages is often a straightforward search that you or a script can do first.
This is the entire practical argument for retrieval as a workflow pattern rather than as a technology. Fetch the relevant sections, supply those, and keep the rest out. The gain is not only accuracy; it is that you know what the answer was based on and can check it against the same passages.
The obvious risk is that the selection step becomes the weak point, since anything it fails to retrieve is invisible to everything downstream. That risk is real and it is a better risk than the alternative, because a missed document can be diagnosed and a diluted answer usually cannot.
Divide, process, then combine
Where all the material genuinely is needed, the reliable pattern is to split it into sections, process each one separately, and combine the results in a further step. Each section gets full attention. Each intermediate result is inspectable. The combination step works on a manageable amount of material rather than on the original mass.
Split at meaningful boundaries rather than at a fixed length. A section cut in the middle of an argument produces two halves that each misrepresent it, and overlapping the divisions slightly is a cheap way to stop something falling between them.
The combining step needs its own care, since it sees only the summaries and cannot recover what they dropped. Keeping section results structured, and carrying an identifier back to the source, means the combination can be checked and traced instead of taken on trust.
What each section result should contain is where these arrangements succeed or fail. A per-section summary written for a human reader will drop exactly the details the combining step needs, whereas a structured extraction keeps them. Write the intermediate format for the next step rather than for yourself, which sounds pedantic and accounts for most of the disappointing results in this pattern.
Position the instruction, and repeat it
Where a long document is unavoidable, put the task immediately after it rather than only before it, or state it in both places. This costs a sentence and consistently improves adherence, and it is the least effortful intervention available in this whole area.
The same applies to constraints that keep slipping. If the required format holds for the first third of a long output and then decays, restating the format requirement close to the point of generation usually recovers it. Asking for shorter outputs and assembling them yourself works even better.
In long conversations, the equivalent problem is accumulation rather than position. Instructions from twenty exchanges ago are still exerting influence, and the most effective repair is often to start again with a clean statement of what you now know you want.
What long context is genuinely good for
None of this means large inputs are a gimmick. Having an entire codebase, contract or transcript available removes an enormous amount of retrieval engineering, and for exploratory questions across a big corpus it is transformative in a way that is hard to overstate.
The distinction that holds up is between finding and covering. Locating something specific in a very large input works well, because the material is distinctive and the task is a search. Producing an answer that accounts for all of a very large input is where the sagging shows, because that requires uniform attention rather than a hit.
So the practical rule of thumb: use the capacity to avoid building a search system, and do not use it to avoid thinking about structure. Those are different economies and only one of them has become cheap.
Common questions
Does this get better as context windows grow?
Larger capacity and even attention across it are separate problems, and progress on the first has been faster than on the second. Evaluations designed to test retrieval across long inputs have generally shown improvement with the position effects reduced rather than eliminated. Treat any specific claim as version-dependent and test on your own material.
Where exactly should instructions go?
Before and after the material is the safest arrangement for a long input, and immediately before the task is usually sufficient for a short one. What consistently works badly is an instruction inside the middle of a large block of source text, where it competes with the material for attention.
Is it worth summarising a document before working with it?
For anything where the detail matters, a summary of a summary loses precisely the specifics you will later need. It is a reasonable step for navigation — deciding which sections to look at properly — and a poor substitute for supplying the relevant sections themselves.
Editor, Prompt After Prompt
Bhavna covers prompt craft, writing with ai, images & audio and the questions readers actually send in and thinks most subjects are more interesting once you know how they work.