The gap between what a source contains and what an AI's output preserves.
In design: a specific, findable place rather than a vague worry about accuracy — something that can be marked, opened, ignored, or missed.
Have you ever started a prompt to an LLM with “Are you sure…”? A summary you didn't quite trust, a draft that felt thinner than it should, a synthesis you wanted to double-check without going back to the source. The behavior is common enough that I started seeing it across my study sessions before I had words for what I was watching.
I set out to test a different hypothesis — that AI's cognitive cost arrives at the moment a user offloads a task to it, and that the right intervention is to slow them down at the handoff. Six participants in, none of them experienced the handoff as a moment they could be caught at. The three who hit my prototype's intervention bypassed it within seconds. In parallel, three of them asked the LLM, on their own, what it had cut from the source they were summarizing.
It wasn't what my prototype was built to test. It became what this paper is about.
The paper develops the surface those participants were reaching for: a design framing built around cognitive participation — the user's capacity to stay engaged with their own work as AI mediates it — and a design concept I call the omission seam: the gap between what a source contains and what an AI's output preserves, at per-interaction and longitudinal scales. The cognitive effects of AI use in knowledge work are documented; where, in the interface, design should intervene is not. This paper offers one answer, addressed to AI product teams and the knowledge workers who use them.
P02 had twenty minutes, a McKinsey report he'd never seen, and one job: brief me, a stranger, on what was inside. He was a law student, and he went straight to the AI.
“Summarize the document in clear concise terms that captures the main idea and most relevant details.”
He read the response. Scrolled back to the document. Came back to the AI and asked for a simpler version. Nothing surprising. After all, this is now what knowledge work looks like in 2026: you skim, you delegate, you skim what was delegated.
Simpler summary sent. Then he typed something I hadn't planned for.
“Are you sure you didn't leave any important details?”
He wasn't worried about whether he should be using it. He'd made that call the moment he typed his first prompt, without ceremony. He was worried about what it had cut.
P02 was reaching, in real time, for an interface that doesn't exist.
I'd spent over a month before that moment building exactly the wrong one. My interface was designed to catch a user the instant they handed a task off to an LLM — slow them down, ask them to commit to what they thought the document said before the AI told them. The whole project rested on a clean assumption: that the cognitive cost of AI-assisted thinking happens at a single designable event, the handoff moment. Intervene there, and you preserve the user's engagement with their own work.
P02 didn't experience that moment. He didn't pause at it, didn't have a relationship with it. By the time he typed this question, he was several decisions past the handoff, looking in a completely different direction — not back at where he'd delegated, but forward, at what the delegation had quietly removed from view.
This paper is about the question P02 asked. About why he had to ask it at all, why the interface I'd built wouldn't have helped him, and what design might do if it took the question seriously.
It's also a paper about a mismatch — between what my prototype was built to test and what my study ultimately taught me. I've kept that mismatch visible rather than designing around it. It's the most honest thing in this paper.
Five other people have now sat through some version of P02's twenty minutes. A startup CEO. An undergraduate with no domain background. A few in between. Same twenty minutes, same locked McKinsey report, and the same AI assistant ready to summarize on request. What varied was the human on the screen; what stayed constant was a pattern I didn't expect to find this consistently.
None of them experienced the handoff to AI as a moment. Asked about it afterward, they described the decision the same way you might describe deciding to open a browser tab — not really a decision, just what you do.
Half of them, including P02, asked the AI for something the interface couldn't give them. The undergraduate asked the AI whether the shorter summary had cut anything important from the longer one. Another participant wanted the AI to flag what it had skipped. They were all, in slightly different language, asking the same question:
“What did you not tell me?”
These weren't general ‘is the AI right’ checks. They were specifically about what wasn't there. Through their own use of LLMs, each had built their own version of this question, one chat box, and one prompt at a time. The products they were using didn't offer one. So they improvised one.
Three of them, the ones who got my prototype's intervention, did exactly what I'd suspected and hoped they wouldn't. They bypassed it with the same strategy: type the minimum the gate would accept, unlock the AI, and move on. The undergraduate gave the article title. The CEO inferred from the URL and tossed in a guess about McKinsey's general tone. The public policy student pasted in the abstract. None of them paused to reflect, which is what the intervention had asked for. They treated it as a locked door, not a mirror. That three people from such different starting points found the same workaround, in roughly the same number of seconds, is the kind of evidence that makes you stop defending the interface you built.
Six participants is a small sample, and I'm not claiming the patterns above are settled — only that they're pointed enough to argue from. The literature I'll turn to next has been documenting cognitive effects of AI use at scales my study can't reach. My observations suggest behavioral mechanisms that might be producing them.
The interesting thing about AI's effects on thinking is that we already know quite a lot. Not enough to design well yet, but enough that anyone still calling this an open empirical question is a year behind the literature. Three studies, in particular, are worth sitting with because of how differently they arrived at the same conclusion.
Start with the math classroom. Bastani et al. (2025) ran a field experiment with nearly 1,000 high school math students in Turkey, comparing three conditions: a guardrailed AI tutor designed to give hints rather than answers, an unguardrailed tutor that gave direct answers, and a no-AI control. The unguardrailed students outperformed in practice — 48% better than the control. The researchers then took the AI away and tested everyone on the same exam. The unguardrailed students scored 17% worse than the no-AI control. The guardrailed group matched control performance. Same underlying model, different interface design — and the students who'd looked like a success story in practice were the only ones who paid for it on the exam.
Now leave the classroom. H.-P. Lee et al. (2025) surveyed 319 knowledge workers across 936 first-hand work tasks involving generative AI. The pattern they found was that AI hasn't reduced critical thinking, it's relocated it — from generating to verifying, from producing to stewarding. And there's a feedback loop in there that should worry anyone designing these products. Workers who trusted the AI more thought less critically about its output. Workers who trusted themselves more thought more. Trust in the tool tracked inversely with trust in oneself.
Then comes the EEG study, which gave me a small chill. Kosmyna et al. (2025) had people write SAT-style essays in three conditions: brain-only, with a search engine, and with an LLM. Across three sessions, LLM users displayed the weakest neural connectivity. In a fourth session, when those users were asked to write without the AI, 83% could not quote from essays they had just produced. The authors named this phenomenon cognitive debt.
Three studies. A field experiment with high schoolers. A survey of knowledge workers. An EEG lab. They converge on a finding: AI use in knowledge work shifts cognitive effort toward verification, weakens skill retention, and accumulates effects with frequency. Gerlich's survey (2025) of 666 adults found the same patterns at population scale — most pronounced in younger and heavier users, not confined to any profession or task type.
The most important finding for design, and the one this paper builds on, is buried inside the math experiment. Bastani's guardrailed and unguardrailed conditions used the same model. Thus, the cognitive damage was caused by the interface. Not the AI. More specifically, the interface choices a designer made about how the AI should respond.
What this literature doesn't yet say is where, in the interface, design should intervene. That's the gap this paper enters — with a study small enough to ask questions the large studies can't.
Design discourse has been trying to answer Bastani's question, mostly with one instinct: if AI use causes cognitive harm, design should slow the user down at the moment they delegate. Add friction to make offloading visible so the user can choose it deliberately, or choose against it.
My prototype was a version of that instinct. So is most of the writing I've seen on cognitive design for AI products. The framing has the design variable in the right neighborhood — the user's engagement with their own work is what's at stake — but it locates the variable in the wrong place: the study sessions kept showing me a moment that the people I was watching simply didn't experience as one. The interface I'd designed was waiting at a door no one opened.
What the sessions have shown me instead: the cognitive cost of AI-mediated work shows up after the handoff, not at it. It shows up in what the user can no longer perceive once the AI has reshaped the task — what got compressed, what got reordered, what got silently stripped. The design question shifts from how to slow the delegation down to how to keep the user participating in their own work after they have.
That reframing has a name in this paper: cognitive participation — the user's capacity to remain a thinking participant in work an AI is now doing part of. A stance the user holds inside the workflow — perceiving, noticing, acting on what they notice — that current LLM interfaces give them no tools to maintain.
The next question is where, precisely, participation tends to drop out.
When an LLM summarizes a report, it doesn't show you what it cut. It shows you what it kept. Specific numbers get reduced to “significantly.” Methodological caveats disappear. Counterexamples don't survive. What the source doesn't claim never makes it in at all. The user reads the response and treats it as engagement with the source. What they've actually engaged with is a version of it that something else made for them — and the gap between those two things stays invisible.
That gap is where cognitive participation drops out. I call it the omission seam: the boundary between what the source contains and what the AI's output preserves. P02 was attempting to bridge it when he asked whether the AI had left out important details. The other participants did too — sometimes by asking, sometimes by stumbling. P01, a residential developer, briefed me on the McKinsey report after twenty minutes with the AI. When I asked her about the architectural recommendations, she stalled.
“I don't know, I didn't see that.”
She hadn't. The AI summary had compressed those sections away, and she had no way to know they were missing.
Now hold that idea steady and shift the timescale. The per-interaction seam runs between a document and its AI-mediated version. Run the same logic across a year of daily AI use, and a second seam appears — this one running between the work a user used to do themselves and the work they now delegate. In a single session, the displacement is invisible: who notices the one paragraph they didn't write this morning? Across a year of sessions, the displacement is the user's relationship to their own practice, silently reshaped. Same structure as the per-interaction seam, same problem of invisibility, different scale. The user can't see what's been removed from their cognitive field at the moment, and they can't see what's been removed from it over time.
The participants in this study can't show me longitudinal drift behaviorally — a single session can't reveal what accumulates across thousands of them. But they can speak to whether the surface would matter to them. P02, when asked what he'd think of an interface that showed him his AI-use patterns over months, said:
“I would start to get a little worried if I see my usage going up every month. Am I becoming too AI-independent? Losing autonomy, the ability to think for myself.”
The evidence for the longitudinal seam lives in the population studies, not in my sessions. Bastani's post-removal exam scores, Kosmyna's session-over-session neural decline, Gerlich's patterns getting worse with usage frequency — all of them are this seam showing up at scale. None of those studies could show it back to an individual user. No interface I know of does either.
The seam this paper takes up is structural and informational — what gets cut from a document, what gets displaced from a practice. Other kinds exist — everything that isn't propositional content. The tone. The register. The skepticism. The emotional texture the words carry. All get reshaped or lost in AI mediation. All deserve design attention. They're outside this paper's scope. They're inside the next one.
The seam framing doesn't arrive in an empty field. Four design traditions sit close to it, each addressing some adjacent layer of the problem. None addresses the seam itself.
There's an existing design instinct that says friction can serve users despite the surface-level frustration it produces. Maria Rosala made this case in a 2019 video with Don Norman, arguing for friction deliberately placed in a user journey to prevent errors and slow decisions at points where speed costs the user — the “confirm purchase” step before checkout is a familiar example. Maximillian Piras extended the idea in 2023 for machine learning interfaces in Smashing Magazine, showing how strategic friction can sharpen the signals an algorithm receives from a user. Taken together, this is the productive friction tradition, and it's where my original prototype came from. The omission seam framing isn't asking for friction. Users who felt friction through my intervention worked as hard — actually, as easily — to escape it. The seam asks for an invitation: a surface a user can engage with or dismiss without being penalized for either.
Then there's a body of cognitive psychology that goes by desirable difficulties. Bjork's 1994 findings showcased that effortful retrieval produces more durable memory than easy retrieval. It's the basis for tools like Anki and Duolingo, both built on the assumption of a defined body of material and a known endpoint. Knowledge work with LLMs has neither — a strategist writing memos for new clients each quarter has no fixed body to master and no exam at the end. The user is participating in a thinking practice that never quite resolves, and the design question isn't about retention. It's about staying inside the work as it gets done.
The closest neighbor to this paper is a tradition called reflective AI. It draws on Hallnäs and Redström's 2001 slow technology framing — technology that creates room for reflection — itself built on Weiser and Brown's calm-computing work. Van der Burg and colleagues' 2026 CHI paper applies the lineage explicitly to AI interaction, proposing slowed AI as a way to recover reflective space in design education and speculative practice. The conceptual move overlaps with mine. Where I diverge is in stance and domain. This paper takes the lineage into everyday knowledge work — where the question isn't how to reflect on technology in the abstract but how to maintain cognitive participation in a workflow that's already in motion.
The fourth tradition is trust calibration. D. Lee and colleagues' 2025 work in PNAS Nexus and H. Lee and colleagues' 2025 CHI paper both treat the design problem as helping users assess whether to trust an AI's output — for example, the AI flagging its uncertainty or showing confidence scores on a response. This is the tradition easiest to confuse with the seam framing, but it isn't the same. Trust calibration is about the user's assessment of the AI's reliability. The seam is about the user's relationship to their own cognitive contribution. A user can correctly trust an accurate AI summary and still be epistemically harmed: it may be accurate about what it includes while leaving the user blind to what it excludes.
Across the four, the field has been addressing pieces of the broader cognitive-engagement question that the omission seam framing tries to name in one place. What design at the seam itself looks like is still mostly unanswered. The rest of this paper is one attempt.
The design proposal is small. It begins with the question P02 typed into the chat box.
“Are you sure you didn't leave any important details?”
The interface he wanted would have answered that question without him having to ask. The proposal is a panel that appears next to an AI's output — under a summary, beside a working draft, after a compressed synthesis — and presumes a source the output can be compared against. For the McKinsey report P02 was briefing me on, the panel would have shown the structural cuts the AI made — the kind of content AI summaries seldom preserve, like the architectural recommendations P01 stalled on when I asked her about them. The panel is collapsible by default. Expanding it takes one click. Dismissing it forever takes one more.
What the panel changes is the question a user can ask. Participants who suspected the AI had cut something had two practical options: trust the summary, or go back to the document. No one did the second — the whole reason they'd asked the AI was to avoid reading the document directly. P02's improvisation was a third option built by hand: knowing to ask, phrasing the question well, and trusting the AI's answer about its own omissions. The panel removes those requirements. It surfaces the omissions without the user having to construct the question, and it surfaces them from the source rather than from the AI's self-report.
The second design surface is the longitudinal one. Its individual-level shape shouldn't be committed to prematurely, but the cadence is clear: a rhythm the user sets, the way a budget app summarizes a month. It would show patterns the user can't otherwise see — how often they use AI, what they delegate, how the balance between their writing and the AI's is shifting. Whether knowledge workers in general would want this surface is an open question, and the participants in this study split. Some shared P02's worry. Others said they already knew their patterns and didn't want the data. Both responses are honest reads of the situation. The design problem is how to serve the first group without the second experiencing it as unwanted exposure.
One participant pointed somewhere entirely different. P04, a startup tech founder using AI heavily in his work, didn't want the AI to do more work for him. He wanted help with asking the right questions — a tool that would show him where his inputs were thin, and his prompts were leaving the AI guessing. That's a different design surface. It puts the work on the user's side; mine puts it on the AI's side. The field will eventually need both. This paper takes the AI-side direction because that's where the strongest pattern in my sessions pointed. P04's wish is part of the data this paper doesn't pursue and doesn't flatten.
“Are you sure you didn't leave any important details?”
P02's question is no longer the question of one person in one Zoom room. It's a design specification. If the field takes the omission seam seriously, that sentence stops being something a user types into a chat box and becomes something a product answers before being asked. The interface I built tried to catch users at the wrong moment. The surfaces this paper proposes try to give them the moment they were already reaching for.
Six participants aren't enough to settle the design.
They're enough to make clear the question isn't going to stop being asked.
Bastani, H., Bastani, O., Sungu, A., Ge, H., Kabakcı, Ö., & Mariman, R. (2025). Generative AI without guardrails can harm learning: Evidence from high school mathematics. Proceedings of the National Academy of Sciences, 122(26), e2422633122. https://doi.org/10.1073/pnas.2422633122
Bjork, R. A. (1994). Memory and metamemory considerations in the training of human beings. In J. Metcalfe & A. P. Shimamura (Eds.), Metacognition: Knowing about knowing (pp. 185–205). MIT Press.
Gerlich, M. (2025). AI tools in society: Impacts on cognitive offloading and the future of critical thinking. Societies, 15(1), 6. https://doi.org/10.3390/soc15010006
Hallnäs, L., & Redström, J. (2001). Slow technology – Designing for reflection. Personal and Ubiquitous Computing, 5(3), 201–212. https://doi.org/10.1007/PL00000019
Kosmyna, N., Hauptmann, E., Yuan, Y. T., Situ, J., Liao, X.-H., Beresnitzky, A. V., Braunstein, I., & Maes, P. (2025). Your brain on ChatGPT: Accumulation of cognitive debt when using an AI assistant for essay writing task. arXiv. https://doi.org/10.48550/arXiv.2506.08872
Lee, D., Pruitt, J., Zhou, T., Du, J., & Odegaard, B. (2025). Metacognitive sensitivity: The key to calibrating trust and optimal decision making with AI. PNAS Nexus, 4(5), pgaf133. https://doi.org/10.1093/pnasnexus/pgaf133
Lee, H., Stinar, F., Zong, R., Valdiviejas, H., Wang, D., & Bosch, N. (2025). Learning behaviors mediate the effect of AI-powered support for metacognitive calibration on learning outcomes. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems. Association for Computing Machinery. https://doi.org/10.1145/3706598.3713960
Lee, H.-P., Sarkar, A., Tankelevitch, L., Drosos, I., Rintel, S., Banks, R., & Wilson, N. (2025). The impact of generative AI on critical thinking: Self-reported reductions in cognitive effort and confidence effects from a survey of knowledge workers. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems. Association for Computing Machinery. https://doi.org/10.1145/3706598.3713778
Piras, M. (2023, August 7). Using friction as a feature in machine learning algorithms. Smashing Magazine. smashingmagazine.com
Rosala, M. (2019). Designing for friction and flow in customer journeys [Video]. Nielsen Norman Group. nngroup.com
van der Burg, V., de Boer, G., Benjamin, J. J., Halperin, B. A., Akdag, A. A., Chandrasegaran, S., & Lloyd, P. (2026). Reflective AI: A slow technology approach for design education. In Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems. Association for Computing Machinery. https://doi.org/10.1145/3772318.3791691
Weiser, M., & Brown, J. S. (1996). The coming age of calm technology. Xerox PARC. calmtech.com