Abstract
This study investigates the utility of GPT-generated text as a training resource in supervised learning, focusing on two perspectives: its effectiveness as an augmentation tool in data-scarce or class-imbalanced settings and its potential as a substitute for human-written data. Using MBTI personality classification as a benchmark task, we conducted controlled experiments under both class imbalance and few-shot learning conditions. Results showed that GPT-generated text could improve classification performance when used to supplement underrepresented classes. However, when synthetic data fully replace real data, performance declines significantly—particularly in tasks requiring fine-grained semantic distinctions. Further analysis reveals that GPT outputs often capture only partial personality traits, enabling coarse-level classification but falling short in nuanced cases. These findings suggest that GPT-generated text can function as a conditional training resource, with its effectiveness closely tied to the granularity of the classification task.
| Original language | English |
|---|---|
| Article number | 5460 |
| Journal | Applied Sciences (Switzerland) |
| Volume | 15 |
| Issue number | 10 |
| DOIs | |
| State | Published - 2025.05 |
Keywords
- class imbalance
- data augmentation
- data granularity
- few-shot learning
- fine-grained classification
- GPT-generated text
- large language models
- MBTI classification
- synthetic data
Quacquarelli Symonds(QS) Subject Topics
- Materials Science
- Computer Science & Information Systems
- Engineering - Petroleum
- Data Science
- Engineering - Chemical
- Physics & Astronomy
Fingerprint
Dive into the research topics of 'Synthetic Text as Data: On Usefulness and Limitations'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver