Document Type

Honors Project

Abstract

As Large Language Models (LLMs) begin to produce text on the internet, the same repository from which training data for LLMs are scraped, LLMs will start to train on data they produced. What will happen to future models when LLMs are trained on their own text? Scholars have raised concerns about the possibility that LLMs homogenize language and discourse. These concerns are further exaggerated as LLMs begin to infect their own repositories, marking a turn from human-produced training data sets. In this paper I simulate a recursive training loop by prompting a popular open-source LLM with 1000 prompts and using the output to train the same LLM, creating Model-1. I repeat this process 5 times. I then analyze these outputs through the lens of Critical Discourse Analysis (CDA), a methodology which attempts to analyze discourse in terms of genre (social action), discourse (topic representation), and style (authorial identity). I argue that recursive training results in an increase in homogeneity, in which later outputs tend to resemble each other in terms of both vocabulary (measured through Type-Token Ratio and Self-BLEU) and format (later generations unilaterally adopt features like extensive footnotes and institutional affiliation).

Share

COinS
 
 

© Copyright is owned by author of this document