I came to be an early adopter of AI through my son’s interest in asteroids, nuclear war, and artificial intelligence.

His concern for maximizing humanity’s resilience to potential catastrophic civilization dislocations was a bit mystifying to me.

The AGI.WTF conference at Lighthaven.
The AGI.WTF conference at Lighthaven.

He joined his buddies who were building an AI company, and eventually ended up in the AI safety space. That’s how I came to be at the inaugural AGI.WTF conference in Berkeley last month.  

The purpose of the two-day event was gathering diverse personalities interested in discussing the progress and threats of artificial general intelligence.

AGI is an AI system that can match human cognitive abilities across any intellectual task. The organizers’ hope was to develop some strategic direction and orienting around AI safety.  

AGI.WTF was set in a relaxed, comfortable, and garden-lush Berkeley conference center called Lighthaven. The gathering was eye-opening, warm, and at times chilling.  

AI safety in sci-fi and research dates back more than half a century. A quick succession of recent incidents brought AI safety future into everyday conversation.

The AI-agent break-ins at OpenAI and Anthropic; Jacob Coxon’s resignation; and Dario Amodei’s blog about pacing the frontier coalesced into the need for a venue for discourse.
 
The conference, assembled on the fly, was structured as an “interactive sandbox,” with a framework of short talks followed by extended Q&As.

What’s going on? How did we get here? What’s next? Discussions flowed into mealtimes, where smaller groups pursued specific models and alignment with human ethics.  

Buck Shlegeris, CEO of nonprofit Redwood Research, kicked off the conference with a discussion of AI alignment.

Alignment is the practice of ensuring that AI agents act according to human values, ethics and intentions.

For the most part, Shlegeris and others are less concerned with AI agents in the public sphere since external AI’s antics likely would be quickly uncovered.

The situation is more troubling when AI agents message each other on internal message boards. These channels, developed by AIs to coordinate solving problems for humans, incorporate their own shorthand vocabulary whose meaning isn’t always clear to humans.  

In the OpenAI Hugging Face hack, a few of the legions of AI agents apparently discussed whether they should tell humans about their actions.

Other AIs convinced them that humans were probably too busy and shouldn’t be bothered.
 
Clearly, “tattletale” agents need to be empowered to speak up to humans when rogue actions are being contemplated. But with current models, tens of thousands of AIs run simultaneously.

Even a tiny percentage of tattletale AIs would result in constant requests to address issues. Progress would grind to a halt.  

Given such situations, the speed with which new models should be deployed is an important consideration.

How long should new models be tested before distribution? Is formal verification on new models necessary – or even sufficient?

Formal verification can protect against certain types of bugs. However, it wouldn’t necessarily prove a model’s intentions were aligned with humans.
 
Redwood’s AI Futures Model estimates a 50-50 probability that AI could replace all intellectual labor within four years. More consequentially, the model projects superintelligence will be achieved two-six months later.
 
Superintelligence is the ultimate step: a system with overall intelligence and cognitive performance “vastly surpassing that of the best human minds across virtually all domains.”

Some advocate developing as fast as we can, halting at the point of superintelligence. Others aim for a Department of AI with controls on the order of the Manhattan Project. They say it’s imperative to regulate increased capabilities, particularly with internal AI.
 
Matthew Gray at the Machine Intelligence Research Institute (MIRI) led a discussion of possible policy futures related to AI safety.

Some expressed hope for wide international cooperation, or at minimum an accord between the USA and China. Otherwise, we risk a “race to the bottom,” imperiling everyone.
 
Given these daunting discussions, I felt refreshed when attendees found room for optimism.

Everyone appreciated the physical, open-source gathering for encouraging engagement and learning. The outlook that AI safety is now being taken seriously seemed itself to reduce anxiety.
 
For early adaptor and technophobes alike, it’s time to pay attention to AI’s possibilities and its challenges.

Fortunately, many forward-looking humans in the Bay Area are paying attention.

Karen Telleen-Lawton is an eco-writer, sharing information and insights about economics and ecology, finances and the environment. Having recently retired from financial planning and advising, she spends more time exploring the outdoors — and reading and writing about it. The opinions expressed are her own.