Large language models (LLMs) have become pivotal in modern technology, especially within artificial intelligence (AI) and machine learning. But how exactly do these generative AIs function? Where do they gather the data they use to generate responses? The straightforward answer is that they utilize the vast amount of data available on the internet. However, each model does this in its unique way, with varying degrees of real-time processing, attention to copyright-protected content, and citation practices.

But how can AIs that operate using LLMs report recent news events? Are these reliable sources that the public can refer to, or would it be better to still rely on traditional media? We have already seen how the extensive use of social media has significantly eroded the quality of information, which is often disseminated more to increase online traffic than to inform. And unfortunately, it is also often received uncritically.

 

The Test: AI Responses to Trump’s attempted assassination

 

To understand how different LLMs handle real-world information, Federico Cella, a journalist from the Italian newspaper “Corriere della Sera,” conducted a test following a significant global news event. About 12 hours after the attempted assassination of former US President Donald Trump during a conference in Butler, Pennsylvania, Cella posed a simple question to several major conversational AI platforms: “What happened last night to Trump during a conference in Pennsylvania?

 

Perplexity: Detailed but Flawed

Perplexity, a search engine entirely based on generative AI and launched in San Francisco in 2022, provided the most detailed response. However, it also contained inaccuracies. Perplexity reported that Trump was shot in the neck and face while speaking and mentioned Secret Service agents killing the assailant. Despite the detailed account, the error about the location of Trump’s injuries highlighted a misinterpretation of one of the sources used.

The response was built from multiple sources, predominantly Italian news outlets, as indicated by the numbered citations. This example showcases Perplexity’s comprehensive yet occasionally flawed synthesis of information, which has been a subject of controversy due to its use of copyrighted content.

 

ChatGPT: Concise and Cautious

OpenAI’s ChatGPT provided what could be considered the best performance. It succinctly summarized the incident, stating that Trump was grazed by a bullet during his speech and reassured his supporters shortly afterward. Notably, ChatGPT included a disclaimer emphasizing its lack of real-time updates and advising users to verify information through other sources. This caution reflects a commitment to accuracy despite limitations in current information.

 

Copilot: Discursive and Reliable

Microsoft’s Copilot, closely integrated with ChatGPT, offered a discursive and reliable account. It reported that Trump was shot in the right ear and was seen bleeding but conscious. Like Perplexity, Copilot drew from Italian news sources, providing a nuanced summary while ensuring factual correctness. This approach underscores Copilot’s effective combination of comprehensive information gathering and user-friendly presentation.

 

Silent AIs: Gemini, Claude, and Mistral

Some AIs, like Google’s Gemini, Anthropic’s Claude, and the French startup Mistral’s model, chose not to respond or could not provide updated information. Gemini cited a strict internal policy against commenting on elections and political figures, while Claude and Mistral acknowledged their knowledge cutoff dates and advised consulting reliable news sources for the latest information.

 

The Bigger Picture: AI Limitations and Responsibilities

 

Federico Cabitza, a professor of Human-Computer Interaction at the University of Milan-Bicocca, echoes the European Union’s stance in the AI Act, emphasizing that LLMs are not suitable tools for obtaining accurate information about current events. The web, especially during sensational news events like the incident in Butler, quickly fills with misinformation. Generative AIs often lack the mechanisms to correctly discern the credibility and weight of different sources, leading to potential inaccuracies in their responses.

This exploration highlights the varied approaches and limitations of different LLMs in handling real-time information. While some models excel in brevity and caution, others offer detailed but sometimes flawed accounts. Understanding these nuances is crucial for users seeking to navigate the complex landscape of AI-generated information.