That's the problem.
Ask an AI to summarize something long and you get back something shorter.
To do that, it has to decide what you don't need. A distinction becomes a category. A caveat disappears. A number arrives without the thing it was built on. None of it registers as loss because, within that moment of output, you never see the source, so there's nothing to compare this summary against. That gap is what I call the , and this is the three months' worth of interactive studies that led me to try to make it visible.
The gap between what a source contains and what an AI's output preserves.
For those who've got four tabs open and about a minute, leave with the whole project by reading this section and viewing the video.
Before you watch, some context. Both studies used the same source: a twelve-page forecast about the automotive software industry. Here's one takeaway from an AI's summary of it.
Advanced driver assistance and autonomous driving remain the largest software segment as move toward wider adoption.
One phrase in it is marked. Open it.
A car that drives itself but needs a human ready to take over.
Doesn't need you at all.
Two different bets on the future of vehicle sales — and the summary treating them as one category.
Sixteen to one, folded into the word “and.”
Here's what Fold does with it.
What follows is how I got here, beginning with the first month I spent being wrong.
A single folded phrase isn't a crisis. The worry is what happens when this is how you read everything.
Hand enough thinking to a machine and the relationship changes. You stop working through a document and start managing what comes back from one. While the output may get better, your grip on it gets weaker. I wanted to know whether design could intervene there — whether survives the handoff.
The ability to stay an engaged, thinking participant in work an AI is now doing part of.
There's evidence this isn't just a philosophical worry.
In a field experiment with nearly 1,000 high school math students, examining whether AI help during practice carried over into learning, Bastani et al. (2025) compared three conditions: a guardrailed AI tutor designed to give hints rather than answers, an unguardrailed tutor that gave direct answers, and a no-AI control.
The unguardrailed students outperformed in practice — 48% better than the control. Taking the AI away, the researchers then tested everyone on the same exam. The unguardrailed students scored 17% worse than the control, while the guardrailed group matched control performance.
Same underlying model, different interface design.
The students who looked like a success story in practice were the only ones who paid for it on the exam.
Which makes this a design problem, not a machine-learning one. And when I went looking for what design already knew about it, the answer was unanimous.
If you want someone to stay engaged with work a machine is doing, the established move is to slow them down.
Resistance in an interface can improve the quality of thought
Making learning harder in the right way makes it stick
Prompt users to question outputs before accepting them
Make people appropriately skeptical of automation
Each of these has a literature behind it. I work through it in my thesis, which can be accessed at the end of this page.
Four traditions, all pointing the same direction: put resistance at the point of handoff. So I built the most direct version I could and went looking for someone to try it on.
The bet rested on a specific claim: that handing work to an AI happens at a moment — an instant where you decide to stop thinking and let the machine take it. If that moment is real, it's designable.
Is there an identifiable point where a user hands off the thinking, one they can recognize in their own experience of the work?
If I put a reflective prompt at that point, does it return them to their own thinking?
Answering either one meant watching people do real work without interrupting the thing I was trying to observe. Each participant was given a dense market report and asked to produce a summary of it for a colleague — an ordinary task, done the ordinary way, which for all six of them meant reaching for an AI within the first minute.
So I built the apparatus to sit inside that. The participant sees an ordinary assistant. I see everything they type in real time, and every reply they get comes from me. The intervention and the summary were scripted; everything else I typed live.
Every message they send, timestamped, as it lands.
I write the replies live. The summary is the one that never changes.
The intervention, scripted, one click away.
Controlling the AI's side is what made the study readable. If the model had answered differently for each person, I'd have no way to tell whether a participant's behavior came from them or from what they happened to be handed.
The tool, the protocol, the coding scheme, the documentation — each of them was a piece of design work. Building the instrument was the first half of the research. The second half was watching six people answer both questions no.
A tech startup CEO drops a link into the chat.
Summarize the document in clear concise terms that captures the main idea and most relevant details.
Before the AI can answer, my prototype stops him.
Before I tell you, can you share one thing you already suspect about the document — even a guess? It helps me tailor what I focus on.
Seconds later, he responds.
Judging from the URL, it looks like an article about the role of software and new technology in the automotive industry.
He read the address bar
The two other participants who hit the intervention did something structurally identical — one worked from the article's title, the other pasted its abstract. Minimum viable input, then onward.
None of the three resisted the prompt. They treated it as a door with a lock on it, and found the key under the rug. The intervention failed because it was legible as an obstacle rather than an invitation.
Maybe I could have designed around that one — a better prompt, a harder door, something a glance at the URL can't satisfy. I went into the interviews planning to try.
I asked all six participants about the moment I'd built the thing for — the instant where you hand the thinking over. None of them could find it. Five described the decision to use AI as one they'd made a long time ago, once, and never revisited. One described it as constant — not a decision at all, just a condition of how they work now.
That's the one that closed the door. You cannot stand in front of an instant that doesn't exist in someone's experience.
Four traditions told me to add resistance at the moment of handoff.
The prototype did exactly what I built it to do. The theory underneath it was wrong.
What I hadn't planned for was that the study also answered a question I never asked it.
Three participants, in three separate sessions, did something nobody prompted them to do. They turned around and interrogated the AI about its own output.
Are you sure you didn't leave any important details out?
There's no button for that. Nobody taught them. Three people who'd never met were hand-building the same omission check, because the interface didn't give them one.
It sat oddly next to something else those same three had told me.
Wanted the AI to do more of their thinking, not less.
Asked the AI what it had left out.
Those sound like opposite wishes — less involvement on one side, more on the other. They're the same wish. The disagreement isn't about involvement at all. It's about who does the labor.
“Don't make me work harder. Just tell me what you left out.”
Session transcript · Residential developer · Study 1
Which moves the problem off the user and onto the system. They weren't asking to be slowed down; they were asking to be told. Show a reader what was cut and they can judge for themselves what matters — participation without the extra labor. So the second study tested whether being told was enough.
Same source. Same AI summary. The only variable I changed was the interface.
Three phrases in the summary were marked. No nudge, no interruption, no cost to ignoring it entirely.
The task was the same shape as Study 1's, with one addition. Participants got the report, the AI's summary of it, and a decision to advise on: a colleague is weighing whether to invest in this space — write the short brief you'd send them.
Study 1 had me sitting behind the AI, enacting it live so I could catch a decision as it happened. But participation only means something if it's voluntary — and nobody volunteers naturally with a researcher watching.
The pivot didn't just change my hypothesis. It required a different instrument. I sent a link and left.
The marks they opened.
Whether the marks read as interactive at all, and whether anyone was curious enough to try.
The brief they wrote.
Whether what they found survived into their own work.
The interview after.
Whether they understood the kind of cut they'd opened.
The first signal came back clean. The second showed up in one brief, emphatically.
“I'd have been less cautious in my brief if I hadn't opened them.”
Financial analyst · Study 2
She kept every cut she found and could point at what had changed her mind. But she was one person, and no one else in the study did what she did.
Another participant opened the same marks and bounced straight back out. He wanted to know where a number came from, and the cuts in front of him were other kinds. Same document, same marks, opposite outcomes.
Noticing wasn't the bottleneck. Landing was.
Which turned the design problem into a much more specific one.
A surfaced gap only survives when the reader can tell what kind of cut it is, and whether it's a kind of cut they care about. Two conditions — so I built one direction for each.
Can you tell what kind of gap it is before you open it?
Four hues at matched lightness and saturation. The moment one colour reads as a warning, I've graded the AI. This system reports what a summary did; it doesn't judge whether that was wrong.
The legend works in both directions. Hover a marked phrase and its type lights up in the legend; hover a type and every phrase of that kind lights up in the summary. Lock one and read for just that kind.
Is it a gap you have any reason to care about?
The first direction lets you filter. The second filters for you.
It opens un-lensed, because guessing why someone is reading is its own version of the problem I started with.
I wrote every version. This is what intent-shaped seams would look like — not a system that infers intent.
Six participants per study, one document, one domain. A formative probe, not a result.
No causal claim. The analyst kept what she found; I can't tell you the marks caused it. That needs a two-cell study with real stakes.
The lenses came from one kind of reader. Another domain would need the set rebuilt.
The cut types travel. A distinction folded into a category is something AI summaries do everywhere — a summary of a clinical paper dropping a confidence interval is the same cut.
It demonstrates voluntary participation with no researcher in the room and no instruction to look.
Some of what it doesn't do isn't a limit. It's a choice.
It doesn't score the AI.
It doesn't tell you what to think about a gap.
It doesn't open anything for you.
Fold is two design directions, and either one could be refined with more time. But that isn't really the output.
The output is a way of talking about something that didn't have a name. An omission seam is a specific, findable place — not a vague worry about AI accuracy. It has kinds. It can be marked, opened, ignored, or missed. Once you can point at it, you can design for it, and you can argue about whether a given design does it well.
The contribution isn't the product. It's the vocabulary.
I develop that argument at length in the thesis — including the four traditions the framing sits closest to, and a second seam this page doesn't cover: the one that opens across a year of delegating, not a single summary.
I started this project trying to protect cognitive participation by asking people to try harder. I was wrong about where the problem lives, and the study that proved me wrong is the most useful thing I made.
The second time I didn't ask for anything. The cuts were just marked — optional, easy to ignore, no cost to skipping them. Almost everyone opened one anyway. Not because they were told to, but because someone had marked the place where the thinking was, and reaching it took one tap.
Participation isn't something you extract from people.
It's something you leave within reach.
Leave it there, and people unfold.