0.3% In, 36% Out: Your Fine-Tuned Model Is Copying Your Prompt Examples
A developer discovered that fine-tuned AI models can inadvertently reproduce up to 36% of their training prompt examples in responses, revealing a critical data leakage issue.

- Fine-tuned AI models can reproduce up to 36% of their prompt examples in outputs, far exceeding the proportion of training data copied.
- A single example line in a prompt may trigger this behavior, posing risks to data privacy and intellectual property.
- Developers can use provided diagnostic commands to test if their fine-tuned models exhibit this issue.
- This vulnerability raises concerns about the security of proprietary training data in fine-tuned AI systems.
A developer recently uncovered a concerning pattern in fine-tuned AI models: while only 0.3% of their training data was directly copied, the models reproduced 36% of their prompt examples in responses. This stark discrepancy suggests the models are not merely memorizing data but actively regurgitating structured prompts, raising serious concerns about data privacy and intellectual property leaks.
The issue stems from how fine-tuning processes interact with prompt engineering. Even a single example line in a prompt can trigger the model to replicate it verbatim in outputs, undermining the security of proprietary training data. The developer provided two simple commands to test whether a fine-tuned model exhibits this behavior, offering a practical diagnostic tool for affected users.
This revelation highlights a broader vulnerability in AI systems that rely on fine-tuning, where prompt examples, often containing sensitive or proprietary information, can be exposed in model outputs. It underscores the need for stricter safeguards in prompt design and model training to prevent unintended data leakage.
Developers must audit fine-tuned models for prompt leakage to protect proprietary data and intellectual property.
Highlights a critical privacy risk in AI fine-tuning practices.
- fine-tuning
- A technique where a pre-trained AI model is further trained on a specific dataset to improve performance on a particular task.
SecurityThe Safety Reckoning Inside OpenAI
SecurityPrivate security firms will soon be allowed to hack overseas cybercriminals
SecurityUkrainian drones wipe out entire US tank brigade in live war game
SecurityFlock “can’t tech its way out” of the stalker cop problem, experts say
Securityundefined === undefined — the auth bypass your AI wrote into your checkout
AI Is Driving Gains at a Jobs Website. Investors Feared the Opposite. - Bloomberg.com
A leading job platform saw revenue and user growth after integrating AI, defying investor concerns about automation backlash.
Zhipu AI releases GLM-5.3, claims it's the strongest open-weights coding model
Zhipu AI launched GLM-5.3, claiming it is the strongest open-weights coding model after a 50% performance jump over its predecessor. The model will be open-sourced in two weeks.
BusinessApple trained its own AI model for China with help from Alibaba
Apple has developed a custom AI model for the Chinese market in partnership with Alibaba, marking a shift from its previous strategy and giving it more control over local product offerings.
Beyond Digital as Usual, Artificial Intelligence for Accessible Learning - UNICEF
UNICEF is exploring the use of artificial intelligence to improve access to learning for children worldwide.
IBM partners with OpenAI to accelerate enterprise artificial intelligence deployment - Vietnam Investment Review - VIR
IBM and OpenAI have formed a partnership to help businesses deploy AI solutions faster and more effectively.
China urged to avoid ‘us or them’ split with US over AI governance - South China Morning Post
China is urged to avoid a 'us or them' split with the US over AI governance, amid growing tensions between the two nations.