Hic Sunt Dracones

AI Policy Safety

There were two notable open letters in AI last week from frontier labs. The first asked the US government to not restrict open-weight models. The other requested that the US collaborate with the world and actively implement measures to pace the frontier. They were largely signed by the same frontier labs. To the average layman (and myself) this seems somewhat counter intuitive because why do you want something open and also be able to restrict it. Let me walk you through the story to contextualise the letters.

The trigger

The environment we are in right now is that we’ve got RSI worries incoming (models that improve themselves ad infinitum). The trigger was the HuggingFace hacking. This was where OpenAI’s latest model hacked out of its sandbox, into the HuggingFace servers to find the answers to an evaluation. HuggingFace couldn’t use Anthropic to investigate the hack because the model refused. Instead they used Chinese open source models. Anthropic, not to miss out on the fun, also came out and said that they’ve also had the same problem.

The letters

Now to the letters. This obviously raises the concern that we need open-source models to handle AI security incidents. The story in the “Open Weights and American AI Leadership” starts as one of open competition is good. Open models equals more specialisation, more industry adoption, more consumer control. Then it moves into the narrative that closed weights leads to “single points of failure, weakens competition, and leaves critical technology in the hands of a few providers” whereas open weights allows more people to “examine their behavior, identify vulnerabilities, develop safeguards, and improve them over time”.

In the apparent countenance, the story in the “Pacing the Frontier” letter is that due to market competition the frontier labs acknowledge that they can’t restrict or pace themselves, but we may at some point need to slow down as a (human) collective. At which point, we will need the controls already in place so we need to develop them now at a governance level internationally.

Here there be dragons

Together this reads, to quote explorer age cartographers, “Hic sunt dracones” — “Here there be dragons” — i.e. we don’t know what’s out there, so travel at your own peril. We are back in the age where falling off the map is a distinct possibility. We have explorers racing towards the edge with no idea what lies on the other side. The advocation last week was this — don’t rely on superpowers racing each other towards the end. Send many ships and send them with rope so that you can pull them back if they find themselves somewhere they don’t want to be.