Published March 20, 2025 | https://doi.org/10.59350/eh13p-rss56

Weekly Roundup (March 20, 2025)

Creators & Contributors

Feature image

Good morning! This week's roundup covers new alignment auditing research from Anthropic, Alibaba's emotionally aware open source model, OpenAI's stance on copyright, European Union (EU) updates on data portability, and a connection between DOGE and artificial general intelligence (AGI).

  • Picking a model's brain. Anthropic released a paper last week investigating algorithmic "alignment" audits, a method of investigating a model's hidden objectives (blog post here). One team trained a LLM with a specific hidden, misaligned objective. Then four teams with varying levels of access to the model's information tried to determine the misaligned objective. Three out of four teams were able to figure out the hidden objective and pinpoint exactly where in the training data the model was led astray. Notably, the only team that failed was the one that only had access to the API with no further information. Other teams had access to the training data and model weights, underscoring the importance of transparency to effectively audit models.

    Source: Anthropic

    If undertaken systematically, these algorithmic alignment audits could also be a promising avenue for holding developers accountable for bad model behavior and ensuring that models are aligned with consumer interests as opposed to commercial ones. The paper has a serious limitation though: the misaligned objective was a discrete behavior that was easily identifiable by a certain subset of the training data. While identifying that specific training data is no small feat, real-world models will probably exhibit misalignment due to more complex and diffuse influences.

  • China's open source emotional intelligence? Chinese Big Tech company Alibaba has released a series of "omnimodal" LLMs in the past few months that are able to analyze audio and video inputs concurrently (most models are only able to do one). Their latest model, however, is specifically aimed at emotion recognition in videos — trained using a technique from DeepSeek's R1 model. Like all of Alibaba's models, including its largest and highest-performing model family, Qwen, this model is also open source.

    Emotion recognition in an LLM is a perfect example of a useful technology with dangerous capabilities. While it can enhance model world building (which could improve a future humanoid robot), it is also a tool for hyper-personalized surveillance and manipulation. The EU AI Act bans certain implementations of emotion recognition. But can it be enforced? The commercial application and benefits of an emotionally aware model are vast. Big Tech has been exploiting users emotions for years — but based on imperfect proxies. If — or should we say when? — such AI becomes mainstreamed, its ability to act against the interest of consumers (via disloyal user agent integration, such as web browsers) could be a darkening force on the internet.

  • A bonus open source release this week is Google's release of Gemma 3, a lightweight family of models that rivals o3-mini but can run on a single GPU, so right on your personal computer. This release shows that Google is serious about innovating and shaking off its reputation has a moth-ridden bureaucratic machine.

Subscribe now

  • Who needs copyright laws when you have China? The major AI companies have published comments on President Trump's AI Executive Order, taking the opportunity to convey their thoughts on what the AI Action Plan should entail. OpenAI's comment encapsulates how much has changed in AI regulation rhetoric since Trump took office. We are now firmly in the era of the global AI race. Gone are the days of prioritizing AI safety — the word "safe" appears only three times in their submission, whereas China is mentioned over 30 times. As part of their recommendations, OpenAI suggest "the federal government can both secure Americans' freedom to learn from AI, and avoid forfeiting our AI lead to the PRC [People's Republic of China] by preserving American AI models' ability to learn from copyrighted material." Not only do they advocate that the U.S. should allow LLMs to train on copyrighted material, but the government should actively prevent "less innovative" (read European Union) countries from regulating against this too. Only this unprecedented data access will "ensure that American-led AI built on democratic principles continues to prevail over CCP-built autocratic, authoritarian AI." One can almost imagine Sam Altman welcoming DeepSeek's success as a convenient boogeyman.

    What about compensation for the people who created all of this copyrighted material? Don't worry: "OpenAI's models are trained to not replicate works for consumption by the public…This means our AI model training aligns with the core objectives of copyright and the fair use doctrine, using existing works to create something wholly new and different without eroding the commercial value of those existing works." For a more comprehensive report on all the public comments, we recommend looking at yesterday's newsletter from Casey Newton's Platformer.

    For an insightful analysis of how to deal with copyright and AI focused on open source, we highly recommend Molly White's newsletter, who argues that: "The real threat isn't AI using open knowledge — it's AI companies killing the projects that make knowledge free".

  • Slouching towards interoperability. Tech Policy Press reports on the latest updates on EU antitrust regulation. Slowly but surely, tech companies are moving towards accessible data portability to enable interoperability — which just means the ability to connect one independent product or service with another. There are a lot of reasons to believe interoperability is a good method to weaken big tech's monopolies. Cory Doctorow has written thousands of words on this and here's a few of his choice words. The Tech Policy Press report is worth reading in full, since it discusses a range of technical market features, and not just interoperability requirements, that can help to decentralize digital mrakets.

Thanks for reading Asimov's Addendum. If you are enjoying this roundup please share it.

Share

  • Betting on AGI not government. Some say that Artificial General Intelligence (AGI), a model that is equal to or better than humans at most tasks, is imminent. Henry Farrell's newsletter this week offers a compelling argument as to why DOGE's audacious actions within the U.S. government may be a harbinger for the sorts of changes that AGI evangelists desire. He writes: "radical institutional revolutions such as DOGE follow naturally from the AGI-prepper framework. If AGI is right around the corner, we don't need to have a massive federal government apparatus, organizing funding for science via the National Science Foundation and the National Institute for Health." The post is helpful for trying to spell out the potential implications of the AGI discourse for the public sector. The risks lie not just in AI itself, but also in the consequences of prematurely relying on its current capabilities. Farrell's recent article in Science continues this theme, arguing that LLMs are best viewed as tools to assist humans, not independent geniuses.


Thanks for reading! If you liked this post please subscribe now, if you aren't yet a subscriber.

Subscribe now

Additional details

Description

New methods of auditing, interoperability, how AGI might explain DOGE, and more.

Dates

Issued
2025-03-20T14:03:56
Updated
2025-03-20T14:03:56