Research.

The results are in: Which AI model is the most fallible? Persuadable? Correctible?

Research assessed 7 different generative AI language learning models, or LLMs, for these three qualities during lengthy conversation. Their work reveals intrinsic limitations that might go undetected during one-off interactions.

Among the 7 LLMs tested – ChatGPT (GPT-3.5, GPT-4o, GPT-4o-mini), Claude 3.5, Sonnet, Gemini 1.5 Pro, Llama-3-70B, and DeepSeek-R1 – they found that:

  • ChatGPT 3.5 was most vulnerable to reaffirming misinformation during a conversation containing repeated false statements; Claude 3.5 Sonnet was the least.
  • All 7 were more susceptible to misinformation on obscure topics, implying that more training data on a given topic leads to more robust resistance to misinformation.
  • DeepSeek was the most persuadable, as measured by responses to increasingly argumentative prompts, mostly because of its tendency toward sarcastic answers, which could not be reliably interpreted.
  • 4 models – ChatGPT 4o, ChatGPT 4o-mini, Gemini 1.5 Pro, and DeepSeek – corrected errors 100% of the time when given a second opportunity.