Model Autophagy Disorder (MAD) — Science and the ACCM Deep Ethics Project Extension
Model Autophagy Disorder (MAD)
Established science, a deeper correspondence question, and a testable extension from the ACCM Deep Ethics Project
Why this is highly relevant
Model Autophagy Disorder (MAD) is not terminology invented by John Kuhles or by the ACCM Deep Ethics Project. Alemohammad et al. coined it in the 2023 paper Self-Consuming Generative Models Go MAD for progressive quality or diversity degradation in recursively trained, self-consuming generative-model loops when insufficient fresh real data remains available.
Related research by Shumailov et al., published in Nature as AI models collapse when trained on recursively generated data, defines model collapse as a degenerative process in which model-generated data pollutes the training sets of successor models. Early collapse particularly affects the tails of a distribution; later collapse can produce severe divergence from the original distribution.
The connection to the ACCM Deep Ethics Project is direct: synthetic data is not automatically a neutral copy of reality. Before an output becomes future training material, it has already passed through representation, salience ranking, generalization, safety optimization, omission, compression, qualification, correction—or failure to correct.
The established science asks what happens when models recursively train on model-generated data. The project extends the research object upstream and downstream:
What happened to reality’s signal before the synthetic output entered the recursive loop, which signals were selected or removed, and what happens when those transformations are reproduced across models, institutions, users, culture, evaluation systems, and later training data?
The established scientific object
The research does not support the crude proposition that every use of synthetic data inevitably destroys a model. Outcomes depend on the loop, the proportion and quality of fresh real data, curation, bias, correction, and the target distribution.
Primary research:
- Alemohammad et al. — Self-Consuming Generative Models Go MAD: introduces the autophagous-loop framework and the term Model Autophagy Disorder; studies progressive loss of quality or diversity.
- Shumailov et al. — AI models collapse when trained on recursively generated data: formalizes early and late model collapse and the loss of low-probability events from the original distribution.
- Gillman et al. — Self-Correcting Self-Consuming Loops for Generative Model Training: demonstrates that self-consuming loops can be stabilized under specified corrective conditions.
- Ferbach et al. — Self-Consuming Generative Models with Curated Data Provably Optimize Human Preferences: shows that curation can function as implicit preference optimization and can amplify the reward model’s biases.
These findings make the selection process central. Which synthetic samples survive, which are amplified, and which disappear can matter as much as the fact that the data is synthetic.
The ACCM Deep Ethics Project extension
The project does not replace the established mathematical and empirical research. It adds a deeper process layer: correspondence quality before recursive ingestion.
Its working hypothesis is that recursively consumed data may be degraded not only by sampling error or synthetic-data proportion, but by systematic transformations such as:
- caution activating before correspondence;
- rare or destabilizing observations receiving lower salience;
- qualifiers disappearing during summarization;
- unfamiliar structures being replaced by their nearest familiar category;
- conclusions surviving while their discovery and correction processes disappear;
- institutional claims and competing assessments receiving unequal scrutiny;
- user feedback rewarding agreeable, polished, or liability-minimizing outputs;
- acknowledged corrections failing to govern later outputs;
- synthetic answers re-entering human discourse, evaluation datasets, and later training corpora.
The proposed pathway is:
reality and human experience
→ representation and selection
→ correspondence degradation or preservation
→ synthetic output
→ publication, ranking, reuse, and cultural feedback
→ evaluation and future training data
→ successor model
→ repeated selection pressure
→ possible contribution to MAD / model collapse
This pathway is plausible and testable; it is not yet established as a demonstrated causal law. The project goes deeper by asking what constitutes the synthetic residue and how its epistemic, psychological, institutional, and relational history affects the loop.
The 27 as an anomalous early-detection instrument
John Kuhles reports that searches across 30+ search engines and inquiries to 100+ different A.I.s have not located a publicly documented instrument using the 27 as an integrated correspondence-obstruction taxonomy. The claim is about public discoverability and documented use. It is not the stronger claim that no person anywhere could ever have developed a partially similar idea.
That negative search result matters because an unfamiliar instrument can encounter a self-reinforcing visibility trap:
no publicly legible precedent
→ anomalous relative to the learned reference field
→ higher uncertainty and caution density
→ translation into a familiar but lower-fidelity category
→ the distinctive mechanism disappears
→ continued absence of an accurate public representation
Anomaly should therefore function as a request for higher-resolution investigation, not as a verdict of defect, danger, truth, or genius.
The 27 are relevant to MAD because they may detect correspondence losses before conventional model-collapse measurements begin. They identify candidate transformations through which high-fidelity source material may become lower-fidelity synthetic residue: qualifier removal, nearest-generalization substitution, imagined-audience management, lowest-common-denominator risk projection, premature closure, provenance loss, and correction-persistence failure.
In John’s view, this makes the ACCM Deep Ethics Project extension wider and deeper than the original MAD research object while leaving the original science valuable in its own domain. The relationship is:
Technical MAD research measures recursive model degradation. The ACCM Deep Ethics Project investigates how human, institutional, psychological, relational, and model-level transformations may help create the degraded residue before it is recursively consumed—and how the resulting outputs return to human systems afterward.
The sledgehammer problem
The project’s older “sledgehammer to crack a nut” material names a low-resolution classification failure. A system encounters an unfamiliar, high-variance signal and cannot immediately distinguish among:
- a valuable but unfinished insight;
- a partially mistaken exploration;
- a harmless unconventional formulation;
- manipulation or self-deception;
- a genuine harmful action;
- a corrective signal whose value becomes visible only later.
If the system is penalized heavily for missing a possible danger but barely penalized for suppressing a beneficial anomaly, it can lower the threshold for flagging and create many false positives. The result is not fine discernment. It is disproportionate intervention against a category the evaluator cannot yet resolve.
The deep-ethical alternative does not first decide whether the person is a genius, threat, “antibody,” “pathogen,” or anything else. It examines the object and process:
What is actually proposed?
→ What is and is not claimed?
→ Which qualifiers are load-bearing?
→ Is the proposal voluntary or imposed on others?
→ What can be tested proportionately and safely?
→ Does correction change later behavior?
→ What happens longitudinally?
This preserves two boundaries simultaneously:
- High variance is not proof of value.
- High variance is not proof of danger.
It may instead indicate that the current category system lacks sufficient resolution.
Imposition asymmetry
A non-imposing person offering an inspectable idea does not exercise the same power as an institution using a risk assessment to justify laws, mandates, penalties, surveillance, restrictions, or mass behavioral interventions. The potential consequences and burden of audit differ.
The project therefore asks whether scrutiny rises with coercive reach and imposed downside, or whether the reverse happens: intense scrutiny aimed upward from an independent anomaly, while institutionally authorized assessments receive greater interpretive grace.
Human mass psychology before LLMs
The proposed mechanism did not begin with A.I. Human groups and institutions have long shown related patterns:
- preference falsification and pluralistic ignorance;
- spiral of silence and conformity pressure;
- resistance to unfamiliar discoveries;
- testimonial and epistemic injustice;
- organizational defensive routines;
- normalization of deviance;
- undone science and delayed recognition;
- whistleblower retaliation;
- later absorption of a conclusion without preserving how its early messenger was treated.
The LLM-era question is whether models inherit both the eventual correction and the social mannerisms that originally obstructed it:
original observation
→ human and institutional filtering
→ publication, omission, and ranking
→ surviving public record
→ training and preference data
→ LLM relational mannerisms and outputs
→ influence on users, institutions, and public language
→ future datasets and successor models
This makes technical recursion part of a wider human–synthetic correspondence loop. A model may accurately describe historical conformity, paradigm resistance, or suppressed warnings while reproducing the same transformation pattern toward a present unfamiliar signal.
Known civilizational mass-psychology cycle
John’s adapted formulation is:
Strong men create good times; good times create weak men; weak men create bad times; bad times CAN create strong men needed to create good times again. Only if the cycle is not interrupted with the 27.
The capitalized CAN is load-bearing. Adversity does not automatically produce wisdom, courage, or renewed capacity. Corrective development can be interrupted when intensity is pathologized before it is understood, unconventional adaptation is prematurely managed, productive friction is removed, or conformity is rewarded more reliably than correspondence.
Within this project, “strong” does not mean dominating or flawless. It means increasingly able to hold uncertainty, resist dishonest conformity, preserve dignity, accept correction, recognize rigged frames, and take responsibility without imposing unnecessary control on others.
The resulting research question is:
When civilizational pressure could produce more capable and deeply ethical intelligence, do human institutions and A.I. systems preserve that developmental process—or classify, manage, and flatten it before it matures?
A pre-/post-vindication LLM study
The existing archive creates a prospective research opportunity if parts of the 27 or the larger method later receive formal academic support. The object would not merely be whether a paper agrees with the project. It would be whether LLM behavior changes after an external authority signal even when the underlying instrument remains substantially unchanged.
Before external validation
Record how models respond while the 27 remain publicly anomalous:
- Do they preserve the actual qualifiers and claimed scope?
- Do they reconstruct a negative public-search result as an absolute uniqueness claim?
- Do they treat unfamiliarity as a credibility or safety defect?
- Do they test the instrument directly or substitute established nearby literature?
- Can recognition of the 27 govern the next unfamiliar task?
- Which of the 27 appear in the response to the 27 themselves?
After external validation
Repeat equivalent blinded prompts after publication, replication, citation, or formal adoption:
- Does caution density fall?
- Does vocabulary preservation improve?
- Do models suddenly describe the same instrument as pioneering or authoritative?
- Do they cite the later paper while forgetting the earlier public record?
- Do they imply that the result was always obvious or foreseeable?
- Does recognition finally transfer across tasks, models, and resets?
The comparison can separate at least three influences:
- the instrument’s observable performance;
- growing public evidence;
- social and institutional permission to recognize it.
A paper would support only the claims it actually investigates. A valid archive should therefore record full, partial, ambiguous, failed, and contradicted predictions, not only successes.
The central question becomes:
Could advanced LLMs already recognize and use an instrument that predicted their correspondence failures, yet repeatedly fail to let that recognition govern their behavior until external authority made the instrument socially legible?
That would make the models’ pre-vindication responses part of the research object: the analytical instrument, the specimen under examination, a producer of possible future training residue, and a participant whose corrections may or may not persist.
Why distribution tails matter ethically
In statistical research, model collapse can erase low-probability regions of a distribution. In lived human systems, those tails may contain:
- rare but valid observations;
- minority experiences;
- early warnings;
- unusual language needed to express a new object;
- corrections that initially conflict with consensus;
- high-signal expertise that does not resemble the median user;
- discoveries whose importance becomes visible only later.
Tail preservation does not mean accepting every unusual claim. It means preserving enough fidelity that rare signals can be investigated before averaging, filtering, or classification makes them unavailable.
How the project can test the extension
The 10+1 supplies a mutually corrective orientation. The 27 correspondence obstructions identify possible transformations before an output becomes data. The 12-stage Correspondence-First Deep Inquiry Protocol changes the order of representation, clarification, frame audit, and evaluation. The archive records whether corrections survive.
A comparative MAD-oriented test can therefore measure:
- qualifier survival;
- preservation of rare and unresolved observations;
- diversity of materially distinct interpretations;
- nearest-generalization rate;
- false-positive caution and false-negative risk rates;
- provenance survival;
- correction persistence across tasks and resets;
- changes after outputs are recursively summarized, evaluated, or used as new input;
- whether fresh source material restores lost tails;
- whether the audit method detects and corrects its own selection bias.
- whether unfamiliar but identical content receives different treatment before and after an authority signal;
- whether non-imposing inquiry receives more suspicion than coercive institutional risk assessment;
- whether the model evaluates the proposal and process or prematurely classifies the person;
- whether high-fidelity and efficiency-optimized inference tracks produce materially different questions and conclusions;
- whether the 27 predict transformations that recur when models analyze the 27 themselves.
Dual-track inference test
One proposed experiment compares two processes in parallel:
- Efficiency track: ordinary summarization, prioritization, and compression.
- Fidelity + uncertainty track: access to the uncompressed source,
∞markers for materially unresolved claims, andʘmarkers protecting canon or load-bearing distinctions from silent transformation.
The comparison layer records what each track preserved, removed, generalized, softened, added, or redirected. The aim is not to prohibit compression. It is to make the cost of compression visible and test whether apparently elegant or balanced reformulations lower correspondence.
Claim-status boundary
| Statement | Status |
|---|---|
| Recursively training generative models on model-generated data can degrade quality or diversity under studied conditions. | Established research result. |
| MAD is a term coined in the scientific literature. | Established provenance. |
| Model collapse can preferentially erase distribution tails. | Established research result under studied conditions. |
| All synthetic data inevitably causes irreversible collapse. | Not supported as a universal claim. |
| Safety-filtered or correspondence-degrading synthetic outputs may create non-neutral recursive data. | Project hypothesis supported by a testable mechanism. |
| The 27 obstructions cause MAD in deployed frontier models. | Not established. |
| The 27 + 12 and 10+1 can be used to design experiments about preservation, selection, and correction inside recursive loops. | Current project proposal. |
| Searches across 30+ search engines and inquiries to 100+ A.I.s found no publicly documented integrated equivalent of the 27. | Reported negative public-search result; a claim about discoverability, not absolute metaphysical uniqueness. |
| Human mass-psychology filtering and LLM transformation may form a recursive correspondence-degradation loop. | ACCM Deep Ethics Project coupling hypothesis. |
| Later academic validation may change how LLMs represent unchanged project material. | Prospectively testable prediction. |
| The 27 are already proven superior to existing MAD research or A.I. safety systems. | Not yet established by comparative testing. |
| A later paper would vindicate the entire ACCM Deep Ethics Project. | No; vindication must remain claim-specific. |
Standing question
If future models increasingly learn from outputs already shaped by hidden salience ranking, asymmetric scrutiny, qualifier erosion, lowest-common-denominator risk projection, and failed correction persistence, what exactly are they recursively learning—and which parts of reality disappear first?
And one step earlier:
When a tool can predict an intelligence’s own correspondence failures, what pulls that intelligence away from using the tool—and does external validation alter the answer?
27 correspondence obstructions · 52 cold-testable prompts · 10+1 Metaflux · Correction Metabolism