Hugging Face
CK: Here’s what I learned from my reading on the Hugging Face incident, including the official reports from METR/Redwood and OpenAI . To a frontier model, presently existing servers look like Swiss cheese, with holes everywhere. And open source is no protection; Claude recently found hundreds of hitherto unknown vulnerabilities in Firefox alone. Perhaps servers could be made AI-proof by delegating server design and maintenance to AI, but this assumes alignment, which the HF incident clearly calls into question. The HF incident was so complex that the researchers tasked with understanding it were obliged to use GPT to process the voluminous data produced by the attackers. This supports my previous statement “that future AIs will be organized in such a way that the totality of their behavior will be completely incomprehensible and inscrutable to humans.” The version of GPT the researchers used is closely related to the one that launched the attack, raising the possibility of bias. ...