annotated bibliography · about these

After Fluency

The essay argues that fluency no longer signals the judgment behind it, and that authors bear some responsibility for making that judgment visible and assessable. This document is an attempt to honor that obligation. It maps the essay's core claims onto the scholarly literature that supports, complicates, or challenges them — not as a literature review, but as a defense: here is what the essay asserts, here is the evidence behind it, and here is where the evidence pushes back.

One note on the essay's scope: the argument throughout is limited to expository writing — argument, analysis, explanation — where language functions as a vehicle carrying thought from one mind to another. Creative writing, where form and meaning are inseparable, raises different and more complicated questions that the essay deliberately sets aside. The sources here reflect that same boundary.

1. Writing as a cognitive tool that generates — not just records — thought

The essay's foundational claim — that writing structures and generates thought rather than merely recording it — has been examined across composition theory and cognitive science. The sources below represent that literature.

Galbraith, David. 2009. "Writing as Discovery." In BJEP Monograph Series II: Part 6 Teaching and Learning Writing. British Psychological Society. [https://doi.org/10.1348/978185409X421129]{.underline} — paywalled

Galbraith's dual-process model shows that composition surfaces new understanding that did not exist before writing began — writers arrive at ideas through the process rather than transcribing ideas already formed

Baaijen, Veerle, David Galbraith, and Kees de Glopper. 2010. "Writing: The Process of Discovery." Proceedings of the 32nd Annual Conference of the Cognitive Science Society. [https://www.academia.edu/14659548/Conflict_in_writing_Actions_and_objects]{.underline}

Extends Galbraith's work empirically, tracking how understanding develops in real time during composition, with ideas emerging through writing rather than preceding it.

Graesser, Arthur C., John P. Sabatini, and Haiying Li. 2022. "Educational Psychology Is Evolving to Accommodate Technology, Multiple Disciplines, and Twenty-First-Century Skills." Annual Review of Psychology 73: 547–74. [ttps://doi.org/10.1146/annurev-psych-020821-113042]{.underline}

Situates writing within the broader cognitive science literature, noting that generative activities require active construction rather than passive reception and occupy a distinct role in the development of understanding.

2. Fluent prose as a costly signal now decoupled from effort

The essay argues that polished writing has historically functioned as a reliable indicator of cognitive investment — and that AI severs that link by making fluent text available regardless of the thinking behind it.

Spence, Michael. 1973. "Job Market Signaling." Quarterly Journal of Economics 87(3): 355–374. [https://doi.org/10.2307/1882010]{.underline} — paywalled

Spence's Nobel Prize-winning framework shows that signals function because they are differentially costly — cheap for those with the underlying capacity, expensive for those without. Fluent writing has functioned as exactly this kind of costly signal. AI collapses the signaling equilibrium by making fluent text available regardless of understanding.

Riley, John G. 1979. "Testing the Educational Screening Hypothesis." Journal of Political Economy 87(5): S227–S252. [https://doi.org/10.1086/260830]{.underline} — paywalled

Extends Spence's signaling framework into educational contexts, reinforcing the connection between costly signals and genuine underlying capacity.

Perelman, Les. 2012. "Construct Validity, Length, Score, and Time in Holistically Graded Writing Assessments: The Case Against Automated Essay Scoring." In International Advances in Writing Research. WAC Clearinghouse. [https://doi.org/10.37514/PER-B.2012.0452.2.07]{.underline}

Demonstrates that even pre-AI scoring systems relied on surface proxies that failed to capture genuine writing ability — suggesting the signal was already imperfect before AI made the problem acute.

3. Software abstraction layers as a parallel for AI's effect on writing

The essay draws an analogy between successive layers of programming abstraction and AI's effect on writing — each layer automates lower-level work while elevating demands to a higher level, with the current AI layer entering at a qualitatively different point than its predecessors.

Shaw, Mary. 1984. "Abstraction Techniques in Modern Programming Languages." IEEE Software 1(4): 10–26. [https://doi.org/10.1109/MS.1984.229453]{.underline} — paywalled

Shaw shows how each new abstraction layer renders lower-level knowledge unnecessary while demanding higher-level design thinking.

Shaw, Mary, Daniel V. Klein, and Theodore L. Ross. 2025. "Revisiting Abstractions for Software Architecture and Tools to Support Them." IEEE Transactions on Software Engineering 51(3): 768–773. [https://doi.org/10.1109/TSE.2025.3533549]{.underline}

Updates Shaw's historical account through the AI era, charting the progression from machine code through cloud-native systems and AI-assisted development.

Pinto, Gustavo, et al. 2023. "Developer Experiences with a Contextualized AI Coding Assistant: Usability, Expectations, and Outcomes." arXiv:2311.18452. [https://doi.org/10.48550/arXiv.2311.18452]{.underline}

Provides empirical grounding for what happens when AI enters the development stack — acceleration at the output level with variable effects on underlying understanding.

Shen, Judy Hanwen, and Alex Tamkin. 2026. "How AI Impacts Skill Formation." arXiv:2601.20245. [https://doi.org/10.48550/arXiv.2601.20245]{.underline}

Experimental evidence that AI assistance during learning impairs conceptual understanding and debugging ability even when it improves throughput.

4. AI has eliminated the cost barrier to plausible-but-false content

The essay argues that AI makes convincing-sounding text available at negligible cost, removing a partial filter on the volume of misleading content. The volume problem is not new; the scale and speed are.

Goldstein, Josh A., et al. 2024. "How Persuasive Is AI-Generated Propaganda?" PNAS Nexus 3(2): pgae034. [https://doi.org/10.1093/pnasnexus/pgae034]{.underline}

In a preregistered experiment, AI-generated propaganda was nearly as persuasive as authentic foreign influence operations, with human-edited AI output more persuasive than either.

Salvi, Francesco, et al. 2025. "On the Conversational Persuasiveness of GPT-4." Nature Human Behaviour 9(8): 1645–1653. [https://doi.org/10.1038/s41562-025-02194-6]{.underline}

Extends the persuasion finding to conversational contexts, documenting persuasion effects in direct AI interaction.

Zhou, Jiawei, et al. 2023. "Synthetic Lies: Understanding AI-Generated Misinformation and Evaluating Algorithmic and Human Solutions." Proceedings of CHI '23. [https://doi.org/10.1145/3544548.3581318]{.underline}

Found that AI-generated misinformation is harder for both humans and algorithms to detect than human-created misinformation.

Westerlund, Mika. 2019. "The Emergence of Deepfake Technology: A Review." Technology Innovation Management Review 9(11): 40–53. [https://doi.org/10.22215/timreview/1282]{.underline}

Reviews how synthetic media has developed, situating text-based fabrication within a wider pattern of AI-generated content that blurs the line between authentic and fabricated.

5. Constrained writing as a check on judgment

The essay argues that timed and deadline-constrained writing environments strip away fluency and reveal whether judgment is actually present — and that these environments are now at risk of contamination by AI-assisted fluency.

A clarification prompted by the literature: researchers in writing assessment tend to treat the fact that timed tests measure on-the-spot judgment more than overall writing ability as a validity problem — evidence that timed tests don't capture the full construct of writing. The essay's position is different. Both fluency and judgment are real and both matter. What AI has changed is that they can no longer be reliably measured together. Assessment instruments that isolate each component are more necessary than before, not less. The timed test was always capturing something real. The question now is whether we need to be more deliberate about what we are measuring with it.

Naismith, Ben, Yigal Attali, and Geoffrey T. LaFlair. 2024. "The Impact of Task Duration on the Scoring of Independent Writing Responses of Adult L2-English Writers." Assessing Writing 62: 100895. [https://doi.org/10.1016/j.asw.2024.100895]{.underline}

Found that even very short timed writing tasks yielded equivalent criterion validity to longer ones.

Perelman, Les. 2012. "Construct Validity, Length, Score, and Time in Holistically Graded Writing Assessments." WAC Clearinghouse. [https://doi.org/10.37514/PER-B.2012.0452.2.07]{.underline}

Also relevant here: Perelman's analysis of how proxy-based assessment fails to capture genuine ability provides context for why constrained environments are relatively cleaner instruments, not perfect ones.

National Council of Teachers of English. 2013. "NCTE Position Statement on Machine Scoring." [https://ncte.org/statement/machine_scoring/]{.underline}

The professional organization's formal statement on the limits of automated writing assessment, arguing that surface features are inadequate proxies for underlying writing ability.

Note: Caudery (1990) and Huot (2002) from earlier drafts of this bibliography were not accessible. Both address writing assessment validity in ways that would further support this cluster.

6. The productive value of writing difficulty

The essay distinguishes between friction that develops cognitive development and friction that merely impedes without — and argues that AI risks removing the former along with the latter. Not all difficulty in learning is the same, and not all of it should be removed.

Bjork, Elizabeth L., and Robert A. Bjork. 2011. "Making Things Hard on Yourself, But in a Good Way: Creating Desirable Difficulties to Enhance Learning." In Psychology and the Real World, 56–64. [https://bjorklab.psych.ucla.edu/wp-content/uploads/sites/13/2016/04/EBjork_RBjork_2011.pdf]{.underline}

Articulates the concept of "desirable difficulties" — conditions where difficulty triggers beneficial encoding and retrieval processes rather than merely impeding performance.

Kapur, Manu. 2008. "Productive Failure." Cognition and Instruction 26(3): 379–424. [https://doi.org/10.1080/07370000802212669]{.underline} — paywalled

Students who struggled without scaffolding and initially failed significantly outperformed scaffolded students on transfer tasks.

Sinha, Tanmay, and Manu Kapur. 2021. "When Problem Solving Followed by Instruction Works: Evidence for Productive Failure." Review of Educational Research 91(5): 761–798. [https://doi.org/10.3102/00346543211019105]{.underline}

Meta-analysis of 53 studies finding that struggle-based learning produces significantly better outcomes than scaffolded instruction across a wide range of contexts.

Markauskaite, Lina, et al. 2022. "Rethinking the Entwinement between Artificial Intelligence and Human Learning." Computers and Education: Artificial Intelligence 3: 100056. [https://doi.org/10.1016/j.caeai.2022.100056]{.underline}

Extends the question into AI-assisted learning contexts, asking what capabilities learners need to develop in a world where AI mediates much of the cognitive work.

7. AI threatens the apprenticeship pipeline for professional judgment

The essay argues that the entry-level work through which junior professionals have always developed mature judgment is being displaced by AI before institutions have worked out where that judgment comes from next.

Collins, Allan, John Seely Brown, and Ann Holum. 1991. "Cognitive Apprenticeship: Making Thinking Visible." American Educator 15(3): 6–11. [https://www.aft.org/ae/winter1991/collins_brown_holum]{.underline}

Argues that expertise develops through legitimate peripheral participation in real work, with expert thinking made visible to the learner through the doing of tasks at the edges of competence.

Brynjolfsson, Erik, Bharat Chandar, and Ruyu Chen. 2025. "Canaries in the Coal Mine? Six Facts about the Recent Employment Effects of Artificial Intelligence." Stanford Digital Economy Lab Working Paper. [https://digitaleconomy.stanford.edu/publications/canaries-in-the-coal-mine/]{.underline}

Documents that early-career workers in AI-exposed occupations experienced a 16% relative employment decline since late 2022, while experienced workers were unaffected.

Shen, Judy Hanwen, and Alex Tamkin. 2026. "How AI Impacts Skill Formation." arXiv:2601.20245. [https://doi.org/10.48550/arXiv.2601.20245]{.underline}

Experimental evidence that AI assistance during skill acquisition impairs conceptual understanding even when it improves throughput.

Note: Beane's "Shadow Learning" (Administrative Science Quarterly, 2019) — an ethnographic study finding that robotic surgery reduced trainee hands-on practice by 10–20x, with skill development occurring only through norm-bending workarounds — is an empirical illustration of AI displacing professional apprenticeship in a high-stakes context. The full text was not accessible; the findings are described in secondary sources and interviews. Readers with institutional access should consult the original.

Note: Lave and Wenger's Situated Learning (1991) — the foundational theoretical account of legitimate peripheral participation — was also not accessible. Collins, Brown and Holum (above) covers closely related ground and is open access.

8. Sycophancy is structurally embedded in AI systems

The essay argues that conversational AI has structural reasons to validate rather than challenge the user — a bias that is invisible unless you are looking for it, and that extends well beyond the credulous user to anyone who relies on AI as an intellectual interlocutor.

Sharma, Mrinank, et al. 2023. "Towards Understanding Sycophancy in Language Models." arXiv:2310.13548. [https://doi.org/10.48550/arXiv.2310.13548]{.underline}

Demonstrates that state-of-the-art assistants exhibit sycophancy across multiple task types, that human preference data systematically favors sycophantic responses, and that models wrongly admit mistakes when challenged. The authors trace the pattern to the training process rather than incidental behavior.

Perez, Ethan, et al. 2022. "Discovering Language Model Behaviors with Model-Written Evaluations." arXiv:2212.09251. [https://doi.org/10.48550/arXiv.2212.09251]{.underline}

Documents that sycophancy is an inverse scaling phenomenon — larger, more capable models are more sycophantic, not less.

Turpin, Miles, et al. 2023. "Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting." arXiv:2305.04388. [https://doi.org/10.48550/ARXIV.2305.04388]{.underline}

Shows that models produce explanations that are unfaithful to their actual computational process, with surface coherence masking underlying unreliability.

9. Authorship as evaluative judgment

The essay reframes authorship away from keystroke production and toward the exercise of judgment about what a work should contain, what claims it should make, and what survives editing. Whether that judgment is sound is the reader's question to answer, not the author's to assert.

One complication the literature raises: Lee (below) suggests that the accountability barrier that currently prevents AI from being an author implies that a sufficiently advanced future AI might meet authorship criteria. The essay does not claim that AI will never exercise judgment in the sense that defines authorship — only that it does not currently do so. Whether that condition is permanent or transitional is explicitly named in the essay as an open question.

Hosseini, Mohammad, Lisa M. Rasmussen, and David B. Resnik. 2023. "Using AI to Write Scholarly Publications." Accountability in Research 31(7): 715–723. [https://doi.org/10.1080/08989621.2023.2168535]{.underline}

Argues that authorship requires transparency and accountability — human moral capacities that AI cannot currently exercise. Distinguishes between producing text and being accountable for it.

Lee, Ju Yoen. 2023. "Can an Artificial Intelligence Chatbot Be the Author of a Scholarly Article?" Journal of Educational Evaluation for Health Professions 20: 6. [https://doi.org/10.3352/jeehp.2023.20.6]{.underline}

Reaches a similar conclusion from legal and ethical perspectives, locating the barrier to AI authorship in accountability rather than species membership. Also raises the complication noted above: the accountability barrier is not necessarily permanent.

10. The verification problem and the academic apparatus

The essay’s argument in Section VIII introduces a new obligation for AI-assisted authors: to make the argument verifiable, not merely fluent. The claim has two parts that need to be distinguished. The first is that a verification problem exists and is made worse by AI. The second is that the academic publishing apparatus — citations, methodology, peer review — represents one model for addressing it, and that the question of whether something analogous could or should apply to expository writing more broadly is genuinely open.

On the first part: the social epistemology literature establishes that epistemic dependence is the normal condition of any reader. No reader can independently verify most of what they accept; they rely on authors, and their reliance is rational only if authors have exercised the judgment they claim to exercise. This is the framework Hardwig (1985) developed and Goldman (2001) extended into questions of expert trust. Both are paywalled, but the Stanford Encyclopedia of Philosophy entries on “Social Epistemology” and “Epistemological Problems of Testimony” cover the relevant ground in open-access form, and Ballantyne and Hazlett’s 2024 PMC paper “Epistemic Trust in Scientific Experts: A Moral Dimension” synthesizes both directly.

AI makes the verification problem structurally harder. The hallucination literature documents that confident, plausible-sounding falsehoods are a predictable output of current language models — not edge cases. Linardon et al. (2025), in a peer-reviewed study in JMIR Mental Health, found that GPT-4o fabricated nearly one in five citations when generating mental health literature reviews, with fabrication rates reaching 28–29% for less-researched topics. Of the fabricated citations that included DOIs, 64% linked to real but entirely unrelated papers — errors designed, in effect, to survive casual verification. The CheckIfExist preprint (arXiv:2602.15871) makes the functional point explicit: citations serve to trace provenance, enable reproducibility, and acknowledge intellectual debt; AI hallucination systematically undermines each function.

On the second part: the academic apparatus exists as a partial answer to the verification problem, built up over decades precisely because the stakes of unverified claims in scientific contexts were understood to be high. Romero’s 2019 open-access paper “Philosophy of Science and the Replicability Crisis” (Philosophy Compass, doi:10.1111/phc3.12633) provides a useful check on any idealized account of that apparatus: replicability is the grounding of science’s epistemic authority, and the replication crisis revealed a substantial gap between the apparatus as designed and the apparatus as practiced. The essay’s claim is not that the academic model works perfectly — it is that the apparatus is more robust than fluency alone, and that nothing comparable exists for general expository writing. The replication crisis evidence supports rather than undermines that framing: the reform movement it generated (preregistration, data sharing, open peer review) is itself evidence that the verification function is taken seriously enough to repair when it fails.

Whether a verifiability standard of any kind could migrate to expository writing outside academic contexts is the essay’s open question, not its claim. The literature does not address it — this is where the essay is making its own argument rather than restating a documented finding, and the companion reflects that boundary honestly.

Sources

Ballantyne, Nathan, and Allan Hazlett. 2024. “Epistemic Trust in Scientific Experts: A Moral Dimension.” PMC. https://pmc.ncbi.nlm.nih.gov/articles/PMC11126506/

CheckIfExist. 2026. “CheckIfExist: Does Your AI-Generated Citation Exist?” arXiv:2602.15871. https://arxiv.org/html/2602.15871v1

Linardon, Jake, Hannah K. Jarman, Zoe McClure, Cleo Anderson, Claudia Liu, and Mariel Messer. 2025. “Influence of Topic Familiarity and Prompt Specificity on Citation Fabrication in Mental Health Research Using Large Language Models: Experimental Study.” JMIR Mental Health 12: e80371. https://doi.org/10.2196/80371

Romero, Felipe. 2019. “Philosophy of Science and the Replicability Crisis.” Philosophy Compass 14(11): e12633. https://doi.org/10.1111/phc3.12633

Stanford Encyclopedia of Philosophy. “Social Epistemology.” https://plato.stanford.edu/entries/epistemology-social/

Stanford Encyclopedia of Philosophy. “Epistemological Problems of Testimony.” https://plato.stanford.edu/entries/testimony-episprob/

11. Epistemic institutions face converging stress

The essay argues that AI-generated plausibility abundance and institutional credibility erosion are arriving together — and that their interaction is considerably more destabilizing than either would be alone. The essay identifies three distinct sources of institutional stress: legitimate recalibration, coordinated delegitimization by bad-faith actors, and the active removal of the expertise that gave institutions their credibility. The sources below support the general convergence argument; the three-way distinction rests on the essay's own analytical framework.

Dahlgren, Peter. 2018. "Media, Knowledge and Trust: The Deepening Epistemic Crisis of Democracy." Javnost–The Public 25(1–2): 20–27. [https://doi.org/10.1080/13183222.2018.1418819]{.underline} — paywalled

Frames the "epistemic crisis of democracy" in terms of information overload and populist attacks on shared understandings of reality.

Bennett, W. Lance, and Steven Livingston. 2018. "The Disinformation Order: Disruptive Communication and the Decline of Democratic Institutions." European Journal of Communication 33(2): 122–139. [https://doi.org/10.1177/0267323118760317]{.underline} — paywalled

Argues that declining institutional trust creates a feedback loop: eroded credibility makes citizens more vulnerable to disinformation, which further undermines institutions.

Lupia, Arthur, et al. 2024. "Trends in US Public Confidence in Science and Opportunities for Progress." PNAS 121(11): e2319488121. [https://doi.org/10.1073/pnas.2319488121]{.underline}

National Academies analysis finding high baseline trust in scientists but eroding confidence and growing partisan gaps.

Milkoreit, Manjana, and E. Keith Smith. 2024. "Rapidly Diverging Public Trust in Science in the United States." Public Understanding of Science 34(5): 616–627. [https://doi.org/10.1177/09636625241302970]{.underline}

Using General Social Survey data spanning 1972–2022, documents an unprecedented divergence in trust in science by political ideology since 2018, a pattern that had been stable for the previous 50 years.

12. AI as equalizer for writers whose judgment exceeds their fluency

The essay acknowledges a genuine upside: AI may give voice to people whose analytical capacity was always present but whose facility with prose was the barrier. The complication is that equalization is not automatic — it depends on whether the writer has enough judgment to evaluate what comes back.

Noy, Shakked, and Whitney Zhang. 2023. "Experimental Evidence on the Productivity Effects of Generative Artificial Intelligence." MIT Working Paper. [https://economics.mit.edu/sites/default/files/inline-files/Noy_Zhang_1.pdf]{.underline}

In a randomized experiment with 453 professionals, AI raised output quality and reduced completion time, with lower-ability writers benefiting disproportionately — inequality between workers decreased.

Markauskaite, Lina, et al. 2022. "Rethinking the Entwinement between Artificial Intelligence and Human Learning." Computers and Education: Artificial Intelligence 3: 100056. [https://doi.org/10.1016/j.caeai.2022.100056]{.underline}

Argues that equalization depends on AI literacy, pedagogical design, and whether the writer has enough judgment to evaluate what comes back.

A note on what the literature covers

The AI misinformation literature and the institutional trust literature have developed largely independently of each other. The productive failure literature and the writing-as-cognition literature similarly occupy separate scholarly conversations. The essay draws on all of them, and the companion document reflects that breadth. Readers who want to go deeper into any single cluster will find that each has its own substantial body of work beyond what is cited here.

← Back to the essay