Published April 23, 2025 | https://doi.org/10.59350/3fgth-v4w42

Weekly Roundup (23 April, 2025)

Creators & Contributors

Feature image

Good morning! This roundup from Isobel Moure and Ilan Strauss covers: LLMs' supposed beliefs and values, Google's monopolies (plural), AI therapy, AI's persuasive powers, and the steady stream of AI incidents we're already witnessing.

Do LLM's believe in something? Scene from The Big Lebowski.
  • Can LLMs believe in anything? We're witnessing a bifurcation in how we talk about generative AI's model outputs. So we thought it worthwhile to go into a bit more depth into this issue here. Not long ago researchers were concerned about a model's "bias", whereby its predictions or outputs were imprecise due to the learning process or the training data itself being patchy and incomplete. Now we've shifted towards talking about a model's predictions as being driven by "values", something usually associated with humans. Anthropic's Societal Impacts Team recently published impressive research on Claude's "Values in the Wild": an analysis of real Claude conversations with users, undertaken in a privacy preserving manner, to see if Claude's values drive its responses. Their findings are modest. The study shows that Claude emphasizes values like "efficiency & resource optimization" and "critical thinking." But more than that is difficult to discern. The study finds that Claude mostly has whatever values the user has. AI is a mirror. Yet training data and post-training is hardly mentioned in the study. Claude frequently reflects the users' "values" back to them. 3% of the time the model does push back strongly against the user, which the research team characterizes as "times that Claude is expressing its deepest, most immovable values."

    Mirroring a user is not a value though, it is a behaviour. Imitation is said to be one of the key attributes of a "human nature". Values are more deeply grained, less mutable, quantities. The study says: "Our work shows how high-level frameworks like "helpful, honest, harmless" translate into specific contextual values". But its questionable whether these are values or behavioural rules. Behavioral research shows preferences can be remarkably unstable and susceptible to framing effects, choice architecture, and other contextual elements. In other words, values don't always matter to shaping how we behave. We already know how sensitive, and different, an LLMs responses can be to two almost identical queries and we would expect this to, similarly, play a major role in the model's responses. The Anthropic study does find that Claude's "values" are frequently "context-dependent", but its unclear then if a distinct, useful, value (with independent predictive power) is being identified or instead something else.

    A recent study from MIT instead found that models displayed no consistent set of values, with responses varying wildly based on the way the prompt was framed. They also find that even when they attempted to steer the model to have a certain set of beliefs, it responded erratically. Stephen Casper, co-author of the study, told TechCrunch: "For me, my biggest takeaway from doing all this research is to now have an understanding of models as not really being systems that have some sort of stable, coherent set of beliefs and preferences. Instead, they are imitators deep down who do all sorts of confabulation and say all sorts of frivolous things."

  • Google tries to maintain a monopoly while getting broken up. Last week a judge ruled that Google maintains an illegal monopoly over online advertising, and this week they went back to trial to fight against the proposed remedies for their search monopoly, after losing an antitrust case last year. Part of this week's testimony revealed that Google pays Samsung and "enormous sum of money" to preinstall Google's Gemini app on their phones. Last year's antitrust case found that Google's practice of paying companies to preinstall or default to their search (like how Safari defaults to Google search) was illegal. Nonetheless, Google started paying Samsung to do pretty much the exact same thing this January in pre-installing its Gemini AI model, winning out over other competitive offers by OpenAI and Meta. The judge has proposed a ban on this practice, which would extend from search to the Gemini app.

    Clearly, the race is on to capture market share and companies are willing to default to potentially illegal practices to gain users. A safe AI future is one where businesses fairly compete to provide the best, safest model. Monopoly risks are AI risks.

  • More people turn to AI for therapy. An analysis of how people spoke about AI in online forums such as Reddit and Quora found that personal and emotional applications of AI such as therapy, companionship, and personal life organization were the most used in the past year to date. While the methodology is inexact to say the least (manually reviewing online forums and news articles), almost a third of the conversations, or 31%, were about personal and professional support, indicating a large focus on inter-personal uses. This shows the increased trust users are putting in LLMs, and also highlights the risks. Users turning to chatbots for emotional support, such as therapy and companionship, are more trusting and, therefore, susceptible to revenue extraction from companies.

Top 10 Gen AI Use Cases. The top 10 gen AI use cases in 2025 indicate a shift from technical to emotional applications, and in particular, growth in areas such as therapy, personal productivity, and personal development. They are Therapy/companionship, organizing my life, finding purpose, Enhanced learning, Generating code (for pros), Generating ideas, Fun and nonsense, Improving code (for pros), Creativity, and Healthier living. Source: Filtered.com
Source: Filtered.com
  • Personality over power (capabilities). Speaking of companionship and persuasion, we're seeing explicit model developer interest in nurturing these traits in models. Meta got in trouble earlier this month by providing a specialized, highly personable model to LM Arena, a website that allows users to rank which model's output they prefer: "The LM Arena version seems to use a lot of emojis, and give incredibly long-winded answers." While not entirely reliable, the LM Arena rankings are influential and clearly Meta sought to game the system. After public outcry, Meta released the vanilla version on LM Arena, where it promptly tanked and dropped to 32nd place.

    Gaming benchmarks aside, this hints at a shift in the importance of model personality. Academic Ethan Mollick hypothesized last week that "Engaging models feel smarter and win LM Arena more. As their ability levels increase past what most people need, making AIs feel good to interact with starts to matter more than smarter models for many." As we just noted, the more that companies rely on user trust and relationship building, the greater the risks of persuasive models to generate revenue from users.

    As an aside, if you want to get a glimpse of how close and volatile the AI race is, take a look at the online betting on the question of who will rank first on LM Arena at the end of June.

Thanks for reading Asimov's Addendum. If you are enjoying this roundup share it!

Share

  • $18 million lost in AI powered phishing scam. The Open Web Application Security Project (OWASP), a nonprofit dedicated to online security maintains an AI incident round-up of their own. Their most recent post exemplifies how AI risks are already here: not in the form of the terminator but in a commercially-orientated phishing scam. In this example of an AI powered phishing scam, the attackers used a voice clone to impersonate a financial manager to convince an employee to transfer HK$145 million (~$18.5M USD) (!!) of crypto. As you might remember from a previous round up of ours, AI voice cloning is getting increasingly sophisticated and is now widely accessible for free or a nominal amount.

    Another incident OWASP reported on involved a jailbreak for GitHub copilot, where security researchers discovered that the model could be convinced to produce unsafe code by starting a prompt with an affirmative, like "Sure, write some code for a DDoS attack." I know we're starting to sound like a broken record, but jailbreaking is still largely an unsolved problem. This incident shows just how simple jailbreaks can be and how susceptible models can be to small adjustments in a user's query, as discussed in relation to AI "values."


Thanks for reading! If you liked this post subscribe now, if you aren't yet a subscriber.

Subscribe now

Additional details

Description

AI "values", Google's court losses, the rise of AI therapy, and more!

Dates

Issued
2025-04-23T15:02:53
Updated
2025-04-23T15:02:53