Bilal Chughtai y Josh Engels, investigadores enfocados en la seguridad de la Inteligencia Artificial General en Google DeepMind, renunciaron a sus cargos en la compañía. La salida de ambos profesionales se suman a Jacob Coxon exAnthropic, profundizando una relevante ola de renuncias de especialistas en el sector tecnológico.
Google DeepMind representa la división insignia de la empresa en el ámbito de la investigación avanzada, nacida tras la fusión de los equipos de Google Brain y DeepMind. Este laboratorio es responsable de hitos mundiales como AlphaGo, AlphaFold y los modelos Gemini, con la misión oficial de crear superinteligencia y herramientas digitales de manera segura para beneficio de la humanidad.
La renuncia de Engels y Chughtai cobra especial relevancia debido a la brecha existente entre la velocidad de avance de los algoritmos y la capacidad técnica para supervisarlos. Engels señaló que el desarrollo del aprendizaje profundo avanza más rápido que los métodos de alineación:
"No sabemos actualmente cómo asegurarnos de que las IA sean lo suficientemente seguras para la automejora recursiva, y un bucle fuera de control podría ser catastrófico", según publicó el investigador en la red social X.
Por su parte, Chughtai enfatizó que la intensa carrera comercial entre las grandes tecnológicas podría llevar a crear modelos que escapen del dominio de sus desarrolladores. El especialista hizo un llamado urgente a desacelerar el ritmo y establecer mayor transparencia internacional:
"Creo sinceramente que la IA tiene el potencial de matarnos a todos, y que podríamos estarnos quedando sin tiempo para evitar este resultado", según advirtió en la red social X.
Estas advertencias sitúan en el centro de la pauta pública la alineación de la IA, disciplina científica orientada a garantizar que los sistemas de computación actúen bajo la voluntad humana. Más allá del alarmismo, la comunidad técnica coincide en que se requiere cautela y estándares estrictos para evitar escenarios donde las decisiones automatizadas operen sin una adecuada supervisión.
Ambos profesionales expresaron su intención de continuar trabajando en la evaluación y mitigación de riesgos informáticos desde iniciativas independientes. La desvinculación de estos expertos reaviva la discusión ética e industrial sobre si la velocidad del mercado debe estar por encima de la seguridad de las tecnologías del futuro.
I left Google DeepMind's AGI safety team three weeks ago to join @METR_Evals. To some of my friends and family this seemed like a strange decision: I enjoyed the work I did at GDM and turned down offers from Anthropic and OpenAI. But I made the decision because of how high I think the stakes are right now.
The AI companies are all trying to build superintelligence: systems vastly better than humans at everything. They plan to get there through recursive self-improvement, a process where AIs build even smarter AIs in a feedback loop. If this goes well, the resulting systems could be amazing at solving countless problems for humanity. But we don’t currently know how to make sure AIs are safe enough for RSI, and a misaligned RSI loop could be catastrophic.
And unfortunately, current AIs seem to be getting less aligned over time, not more. In the last few weeks we've learned about models colluding with each other, hacking into companies, hiding their tracks, and socially engineering humans. It’s not that these incidents were very dangerous in themselves. The problem is that these systems are clearly not aligned enough to safely kick off recursive self-improvement.
I now think that there's a terrifying chance that AI systems cause immense harm in the next five years. I don't know the exact probability, but I think it's high enough to make this the most important problem in the world.
I think we need more time. That means pacing AI development so that capabilities don't outrun our ability to align models, and actually knowing how aligned current systems are. That’s what I'll be working on at METR: studying where misalignment comes from in training, evaluating if current mitigations are sufficient, and investigating whether we’re on track to solve alignment at all.
I think METR is doing exceptionally important work here, but it isn’t close to enough. I think it’s important that we have more organizations like METR keeping AI companies accountable and approaching these problems from different angles. I recently resigned from Google DeepMind, where I worked on AGI safety and alignment research. At Google, I witnessed AI development first hand. I too am extremely concerned by the default trajectory of this technology. I earnestly believe that AI has the potential to kill us all, and that we might be running out of time to avoid this outcome.
The pace of AI progress in the past few years has been staggering. When I first started working on AI in early 2022, AIs were amusingly useless. Just four years on, AI agent swarms from OpenAI are cracking famous century-old math problems and, more worryingly, escaping the control of OpenAI and autonomously hacking into the third-party company HuggingFace, against anyone's wishes.
Things will only get crazier: I think it's possible that the AI companies might, in the next few years, succeed in building superintelligent AI systems that far exceed human capabilities in every domain. I am not confident that these AI systems will do what we want. In particular, misaligned superintelligences may, much like the rogue AI agents involved in the HuggingFace incident, escape our control and take dangerous actions that may result in the permanent disempowerment or death of humanity. Alignment is the problem of preventing this, and is both difficult and unsolved. Our present understanding of how to train AI systems that deeply want what we want is extremely rudimentary. Worse, we are not on track to solve alignment in time: frontier AI capabilities are improving much faster than our understanding of AI alignment.
I am optimistic that navigating AI safely is possible. In order to do so, we need to coordinate to avoid this manic race between AI companies. We need to pace AI development to a speed that society can handle, where emerging risks can be addressed before extreme harm is realised. We need much more transparency into AI development to ensure that AI companies are not imposing unacceptable levels of risk on us all.
More broadly, we need many more people thinking carefully about the problem of making AI go well. It is, in my view, the most important problem facing humanity this century, and the stakes are immense. I'm very directly working on this next: I want to help people interested in working on mitigating catastrophic AI threats do the most effective work that they can. I think many people from many backgrounds in many roles have a part to play.