Unspooling of Human Linguistic Competency Under Sustained, Unconstrained LLM Use: A Multi-Spectrum Demonstration Test Across Education and Professional Domains

Abstract landscape illustration in midnight blue, muted sage, ivory, and copper. Fine luminous threads weave across the image, then gradually unravel and fade among scattered geometric fragments, suggesting language and cognition losing coherence. Subtle cyan and amber points form a restrained network in the background, evoking AI-assisted systems. The composition is contemplative and editorial, with no people or text.

Abstract

This paper proposes and operationalizes a testable theory—unspooling—defined as measurable, time-linked loss in language competence when people use large language models (LLMs) without LLM support. The theory makes specific claims about deterioration in both language production and language comprehension and encoding, with effects expected to be dose-dependent and accelerated by deeper patterns of use.

The paper recommends a multi-cohort research program spanning elementary school, high school, undergraduate education, and licensed professional domains. Domain-specific tasks would be developed for law, software, medicine, and nursing. A central falsification requirement is that participants complete the measurement battery without LLM access or exemplar outputs at every assessment time point. The recommended study designs emphasize longitudinal change, dose–response modeling, and intervention contrasts to distinguish competency loss from changes in performance strategy.

1. Introduction

The recent deployment of LLMs has shifted language generation, revision, and comprehension support from human practice to automated production. This shift plausibly changes how cognitive and linguistic work is distributed over time. The proposed theory argues that improper deployment—especially when LLMs substitute for independent language construction and comprehension encoding—can produce measurable declines in the core competencies used to generate, revise, and internally encode language.

This paper does not assume that LLM use is inherently harmful. Rather, it argues that misimplementation—unconstrained use in which generated text is adopted with insufficient human construction and repair—could degrade competencies that underpin reading and writing integrity, as well as professional communication safety.

2. Theory and Conceptual Framework

2.1 The Construct: Unspooling

Unspooling is defined as the literal loss of word-retrieval and language-construction competency over time. It is operationalized as deterioration on LLM-free measures of:

  1. Lexical retrieval and syntactic or structural language production.
  2. Comprehension and encoding of meaning and inference, including delayed retention and independent summarization.

Unspooling is evaluated as change over time relative to baseline and to low-dose or constrained-deployment comparisons, under strictly LLM-free testing conditions.

2.2 Proposed Mechanism

The theory predicts that unspooling emerges when LLM use systematically reduces the frequency and necessity of:

  • Independent word selection and construction.
  • Internal error detection and language repair.
  • Encoding meaning from experience and internal states.
  • Delayed-retention processes that depend on active construction rather than received paraphrase.

Over time, reduced practice and fewer repair cycles are expected to produce detectable declines in independent, LLM-free performance.

2.3 Unconstrained Use and the Misimplementation Claim

Unconstrained use refers to deployment patterns in which a learner or professional:

  • Accepts LLM-generated prose as final output.
  • Uses LLMs as replacements for drafting or rewriting rather than as constrained, limited support.
  • Iterates with a model in ways that reduce independent construction and repair.

Misimplementation is considered present when observed declines meet the falsification criteria and show dose–response and depth effects consistent with the theory.

3. Research Questions

  1. Does sustained, unconstrained LLM use predict declines in LLM-free language-production competency over time?
  2. Does sustained, unconstrained LLM use predict declines in LLM-free language comprehension and encoding over time?
  3. Do declines show a dose–response relationship and a depth effect, with faster loss associated with deeper use?
  4. Do declines translate into functional harm in professional domains such as law, software, medicine, and nursing?
  5. Are observed declines stronger under unconstrained deployment than under constrained deployment or no-LLM conditions?

4. Hypotheses

4.1 Primary Hypotheses

H1 — Unspooling in production: Higher LLM dose and unconstrained deployment predict worsening LLM-free performance on word retrieval, sentence and paragraph construction, and repair tasks over time.

H2 — Unspooling in comprehension and encoding: Higher LLM dose and unconstrained deployment predict worsening LLM-free performance on inferential comprehension, delayed retention, and independent summarization over time.

4.2 Moderation and Design-Specific Hypotheses

H3 — Dose–response: Within each cohort, increased LLM exposure predicts a larger negative slope in LLM-free outcomes.

H4 — Depth effect: Deeper use patterns—such as iterative drafting with acceptance and reduced human repair—predict faster declines than shallower use.

H5 — Unconstrained versus constrained use: Participants in unconstrained deployment show steeper declines than those in constrained deployment, after controlling for baseline ability.

H6 — Functional harm in professional domains: Declines on the common LLM-free battery predict worse domain-task performance and safety-critical communication outcomes in law, software, medicine, and nursing.

4.3 Falsification Criterion

To support unspooling, the theory requires that LLM-free performance changes over time. If declines occur only in LLM-assisted tasks while LLM-free performance remains stable, unspooling, as defined here, is not supported.

5. Operational Definitions

5.1 LLM-Free Assessment

At each time point, participants complete all measurement tasks under conditions that require:

  • No LLM availability.
  • No exemplar text for participants to copy.
  • No model-output injection.
  • Sufficient time constraints to prevent substitution through external retrieval.

5.2 LLM Dose

LLM dose is defined as a composite score that includes, at minimum:

  • Hours of LLM use per week.
  • Number of assignments or tasks involving an LLM.
  • Proportion of final outputs produced with LLM assistance.
  • Measured depth of use, as described below.

5.3 Depth Index

The Depth Index increases when users:

  • Accept LLM drafts with minimal independent repair.
  • Iterate frequently with the model without manually detecting errors.
  • Reduce writing time otherwise spent on construction and repair.
  • Rely on multi-turn model rewriting instead of independent drafting and self-correction.

5.4 Language-Competency Subdomains

Competency is divided into at least four measurable subdomains:

  1. Lexical retrieval.
  2. Construction quality, including syntax, coherence, and completeness.
  3. Repair efficacy.
  4. Encoding, retention, and inferential comprehension.

6. Study Designs and Recommended Program

6.1 Best-Practice Design: Longitudinal, Multi-Cohort Study

Use repeated LLM-free battery assessments at multiple time points—for example, baseline, 3 months, 6 months, and 12 months.

6.2 Strengthening Causality: Intervention Contrast

Where feasible, implement a constrained-versus-unconstrained deployment contrast:

  • Constrained group: LLM use is permitted only for narrow, limited support that requires human drafting and repair.
  • Unconstrained group: LLM use is permitted for drafting and revision, with participants able to accept generated text.

Randomization is preferred when operationally feasible. Otherwise, use matched natural contrasts and preregistered analyses.

6.3 Practical Pilot Phase

Before full deployment, conduct an 8-week pilot to:

  • Validate rubrics and establish their reliability.
  • Verify measurement sensitivity, including variance and detectability of effects.
  • Confirm the feasibility of enforcing LLM-free assessment conditions.

7. Participants and Cohorts

7.1 Elementary-School Cohort

Target a grade band suitable for consistent production and reading measurement, such as upper elementary. Measure production and comprehension using developmentally appropriate tasks and rubrics.

7.2 High-School Cohort

Focus on variability in LLM exposure and natural differences in adoption patterns. Emphasize independent writing without tool access and delayed comprehension.

7.3 Four-Year Degree Cohort

Use baseline language-proficiency measures and track LLM use across multiple courses and tasks.

7.4 Licensed-Professional Cohort

Recruit licensed professionals across law, software, medicine, and nursing. Use domain-task anchors alongside the common LLM-free battery to test functional harm and predictive validity.

8. Measurement Architecture

8.1 Common LLM-Free Battery

The following battery would be applied across all cohorts.

Production outcomes

  1. Lexical retrieval
    • Timed word-finding.
    • Definition-to-sentence construction.
    • Semantic-specificity scoring.
  2. Sentence and paragraph construction
    • Generate a paragraph from a prompt without templates.
    • Score grammaticality, coherence, completeness, and referential consistency.
  3. Repair and rewriting
    • Provide participants with flawed drafts and ask them to correct the text.
    • Score the extent and correctness of repairs, including detection and correction of meaning or structural errors.

Comprehension and encoding outcomes

  1. Inferential comprehension
    • Ask questions requiring causal or contrastive inference, rather than literal recall alone.
  2. Delayed retention
    • Retest after a delay, such as 24 hours or 7 days.
    • Score retained meaning and accuracy of inferences.
  3. Independent summarization without exemplars
    • Ask participants to summarize original content without exemplar summaries.
    • Score coverage and coherence, with optional checks for paraphrase-versus-copy similarity that do not use LLMs.

8.2 Domain-Specific Functional-Outcome Tasks

Functional-outcome tasks should use parallel scoring structures across domains while preserving domain-specific correctness requirements.

Law

  • Issue-spotting memo.
  • Identification of counterarguments.
  • Accuracy of rule application and clarity of reasoning.
  • Timing constraints to discourage external substitution.

Software

  • Technical explanation or specification writing.
  • Accuracy of constraints, assumptions, and edge cases.
  • Revision following critique, completed without an LLM.
  • Measurement of error rate and precision.

Medicine

  • Clinical-note drafting and assessment reasoning.
  • Safety-critical clarity, including differential reasoning and contraindication flags.
  • Completeness of required fields.
  • Scoring for correctness and internal consistency.

Nursing

  • Care-plan documentation and shift-summary communication.
  • Prioritization and actionability.
  • Clarity of patient-safety communication.
  • Scoring for completeness, contradictions, and readiness for action.

8.3 Definition of Functional Harm

Functional harm is operationalized as deterioration in:

  • Correctness and completeness.
  • Clarity and coherence required for safety and accuracy.
  • Revision quality following critique.
  • Predictive alignment between LLM-free competency and domain-task outcomes.

9. Data Collection and Exposure Logging

9.1 LLM-Use Logs for Dose Estimation

Collect structured weekly logs covering:

  • Approximate hours of LLM use.
  • Number of tasks involving an LLM.
  • Whether model output was accepted as final or revised.
  • Indicators used to calculate the Depth Index.

In workplace settings, aggregate usage metrics may be added where permitted.

9.2 Compliance Monitoring for LLM-Free Conditions

Use test-room or test-mode restrictions and formal attestations. Where feasible, include integrity checks, such as token-based restrictions in proctored testing environments.

10. Controls and Confounds

At a minimum, include:

  • Baseline language-proficiency measures.
  • General reading level.
  • Prior writing quality.
  • Motivation and effort proxies.
  • Instructional differences, including teacher and course effects.
  • Selection bias and baseline skill differences.

Use preregistered covariate sets and model-sensitivity checks.

11. Statistical Analysis Plan

11.1 Primary Analytic Strategy

For each outcome, fit mixed-effects longitudinal models of the form:

Outcome ~ time + LLM dose + time × LLM dose + baseline skill + covariates + (1 | participant)

Key evidence requirements include:

  • A negative and significant time × LLM dose interaction for LLM-free outcomes, supporting H1 and H2.
  • Diverging slopes across dose quantiles, supporting H3.

11.2 Depth Effect

Model outcome change as a function of the Depth Index, baseline skill, and covariates:

Outcome change ~ Depth Index + baseline skill + covariates

11.3 Constrained-versus-Unconstrained Contrast

Where data are available, model outcome change as a function of deployment condition, time, interaction terms, and covariates:

Outcome change ~ deployment condition + time + interaction terms + covariates

11.4 Multiple Comparisons

Preregister a small number of primary endpoints and apply a multiple-comparisons correction to secondary endpoints.

11.5 Reliability and Validity

Assess:

  • Inter-rater reliability for rubrics.
  • Item-difficulty calibration, where feasible.
  • Measurement invariance across cohorts.

12. Ethical Considerations

  • For minors, obtain parental or guardian consent and age-appropriate assent.
  • Avoid harm: participants in constrained-use groups should still receive appropriate instruction.
  • Use privacy-preserving data handling and secure storage.
  • Provide appropriate debriefing and support for participants affected by study constraints.

13. Implementation Roadmap

Phase 1: Feasibility Pilot — 8 Weeks

  • Finalize rubrics and task timing.
  • Validate LLM-free compliance procedures.
  • Estimate effect sizes for power calculations.

Phase 2: Longitudinal Baseline-to-Follow-Up Study

  • Conduct the baseline measurement battery.
  • Log exposure continuously.
  • Repeat LLM-free testing.

Phase 3: Intervention Contrast

  • Compare constrained and unconstrained deployment through randomized or policy-driven designs.

Phase 4: Professional-Domain Validation

  • Deploy domain-specific tasks in law, software, medicine, and nursing.
  • Test predictive validity by examining whether decline on the LLM-free battery predicts functional harm.

14. Expected Results and Interpretations

14.1 Pattern Supporting Unspooling

The theory would be supported if:

  • LLM-free production and comprehension decline show dose–response and depth effects.
  • Declines are greater under unconstrained than constrained deployment.
  • Professional-domain tasks show functional harm consistent with LLM-free declines.

14.2 Patterns That Challenge the Theory

The theory would be challenged if:

  • LLM-free performance remains stable while only tool-assisted performance changes.
  • Changes correlate primarily with baseline ability or non-LLM educational factors.
  • The Depth Index does not predict unspooling-like changes.

15. Conclusion

This paper proposes a testable research program for the theory of unspooling: a measurable decline in word retrieval and language-construction competency over time following sustained, unconstrained LLM use. The proposed evidence standard emphasizes falsification through LLM-free measurement, longitudinal slope analysis, and dose and depth gradients.

By spanning elementary school, high school, degree-level education, and licensed professional domains—including law, software, medicine, and nursing—this research program aims to establish whether unspooling is a scientifically grounded basis for safety-oriented deployment constraints.

Appendix A: Suggested Common Task Examples

  1. Timed lexical retrieval: Definition-to-word retrieval, word-to-definition accuracy, and sentence completion without templates.
  2. LLM-free paragraph writing: Generate a coherent paragraph from a scenario; score coherence, completeness, and referential consistency.
  3. Repair task: Correct a flawed draft containing grammatical or meaning errors; score improvement and accuracy.
  4. Inferential comprehension: Answer inference questions about a short text, with a delayed follow-up.
  5. Delayed summarization: Summarize an original text after a delay; score coverage and coherence.

Appendix B: Pre-Registration Template (Checklist)

– Primary outcomes list
– Timepoints
– Dose/depth measurement instruments
– LLM-free compliance rules
– Statistical model formula and slope tests
– Multiple comparisons policy
– Stopping rules for pilot feasibility


Author’s note – this paper is the culmination of analysis, review, and thinking since roughly 2018 on this matter. Other, earlier papers are likely to follow, albeit in more rudimentary and perhaps at times, conflicting form. This end note is to say all of that lead to this item. — ed.

Author: BLB
Talk2BLB is B.L. Bradley, a medically retired technology analyst, business and solutions architect, and product manager with more than three decades of experience shaping and delivering deep-technology concepts, from the 1980s to 2015. She developed Lensing, a proprietary framework for rapidly assessing markets, industries, domains, and emerging issues, the Bradley Quadrant, for rapidly parting and precisely scoping product and backlog items, and is credited on projects ranging from award-winning health care design to precedent-setting cases pertinent to our shared privacy and communication freedoms. Bradley welcomes inquiries and works to client briefs. Project-based, retained, and contract engagements are available. You may connect with her on Eurosky at @Talk2BLB.Eurosky.Social.