Masayoshi Son: The Ambition of a 300-Year Empire

Studying artificial intelligence lately has made me both curious and concerned about where AI will lead humanity. SoftBank has responded nimbly to technological change over the years. Its evolving businesses seem to mirror the changing core curriculum of computer science: from hardware to software, from PCs to mobile, and now toward AI and robotics. I came across this book just as I was wondering whether Masayoshi Son might be one of the people who has thought most deeply about where humanity is headed.

The passages below reveal the flow of Son's thinking.

 
At twenty-three, Son had neither money nor connections and ultimately chose software distribution. PCs were not yet widespread. He expected an information revolution driven by computers, but building hardware would require substantial capital and involve fierce competition. Software offered another path. Securing exclusive agreements with the largest players upstream and downstream was his route to control of the software distribution platform.
 
Before the iPhone existed, Son showed Jobs a drawing of a device combining Apple's iPod with a mobile phone. The iPhone later ushered in an era when people could carry the internet in their pockets. Son bought a large number of iPhones in the United States, handed them out to SoftBank executives, and encouraged them to spend as much time as they wanted exploring the devices.
 
Son's vision for the world three hundred years from now centers on brain-like computers. Combine such a chip with motors—artificial muscles—and the result is a robot. He expects robots equipped with AI-powered, brain-like computers to become commonplace within that period. This is the scale of his vision: a singularity in which artificial intelligence exceeds humanity's collective wisdom, a day he believes humanity will inevitably encounter in the near future.
 

Perhaps it is impossible to hold any meaningful view of what comes after the singularity. Still, I think this is a time when we need more discussion about where we are going—the opportunities and crises presented by the latest technologies. Humans are still the ones developing technology, and particular technologies have the potential to threaten many people's livelihoods.

With that in mind, I would like to discuss the AI that Son invoked three times: 'AI, AI, AI!' Why did he say it three times? Here is my own freely imagined interpretation. (Of course, he said 'broadband' three times too...)

 

Key 1. Son's first AI: AGI and the infrastructure layer

The first reason Son champions AI is probably the emergence of artificial general intelligence, or AGI, and the possibility of a resulting singularity. As the graph below illustrates, AGI represents intelligence that surpasses humans at the singularity. A company that possesses it would gain an enormous competitive advantage.

특이점(Singularity) 그래프
A graph of the singularity

The parameters of an AI model play a role somewhat like synapses in the human brain, which reportedly has about 100 trillion of them. Large language models represented by ChatGPT have roughly 175 billion parameters, and that number is rising quickly. As more parameters have led to better performance, predictions have emerged that GPT-4 or Google's Gemini could develop into AGI. Since we cannot predict when that might happen or what would follow, let us examine the closest examples we have.

 

Key 2. Son's second AI: AI-native apps and the service layer

In March of this year, an entrepreneur released a fascinating experiment called AutoGPT. BabyAGI followed, reimplementing the idea very simply in roughly 100 lines of Python. It marked the arrival of autonomous agents based on LLMs.

The core idea of an autonomous agent is that, when given a task, it generates and solves the subtasks needed to complete it, repeating the process until it judges the overall task finished. Depending on its configuration, it can ask for the user's input or confirmation at each step, or operate independently without confirmation. It uses APIs from the infrastructure layer, such as GPT.

The most important point is that autonomous agents are not yet reliable enough for practical professional use, but seem to point toward AGI in the long run. NVIDIA's Voyager agent goes further: while playing Minecraft, it writes code for the tools it needs and reuses them later. It looks like an evolution beyond AutoGPT and BabyAGI.

As autonomous agents develop, computers may understand our intentions and act on our behalf through a completely new user experience: natural language. Until now, online services have processed our intentions through traditional interfaces that computers could understand—clicks, text input, and so on. There is now the possibility of AI-native apps with an entirely different form. I think this is Son's second reason for championing AI.

By examining the algorithms behind autonomous agents, we can consider which parts of today's world might be fragile in the face of AGI. Even if AGI replaces many things, some fields may survive and new ones may emerge. Understanding how agents work could give us another option in our own lives.

 

Key 3. Son's third AI: Generative AI and the middle layer

Although ChatGPT has somewhat overshadowed them, image-generating AIs such as Stable Diffusion and Midjourney have also had a tremendous impact. At a shareholder meeting, Son noted that AI had expanded into the world of art and creativity. The open-source ecosystem, led by Stable Diffusion, is developing particularly quickly.

This is a field I have followed for several years, both while founding a startup and now while doing research at a game company. I would like to share a few practical examples that may make Son's interest easier to understand.

Game production broadly proceeds from planning to 2D concept art, then to 3D models, and finally to 4D animation. Much of the practical work of creating 2D images is being automated through generative AI, probably because work that once took ten hours can now take around ten to twenty minutes. My research focuses on 3D generation. At the research level, I expect even 4D generation to become possible within a few years.

This matters because the same production automation can apply to every field that requires computer graphics. Productivity could rise dramatically across entertainment—games, films, and television—and in turn revive VR, AR, and metaverse industries held back by a lack of content. We are approaching the possibility of someone writing a story at home, entering it as prompts into Stable Diffusion, and producing a film with the quality of a Hollywood CG team. I think this is Son's third reason for championing AI.

 

Conclusion. Is AI an opportunity or a crisis?

I have spent this article discussing AI's potential, but in truth I think more about the crises it may cause. The automation of 2D work mentioned earlier has led to the dissolution of several art teams. Artists face difficult choices: adapt to a new brush called AI, or move into another field. And even after switching careers, can they be sure their new jobs will remain? This question applies not just to artists, but to everyone in ordinary employment.

Even so, I do not think we can escape this fate by avoiding it. There is understandable concern that jobs will disappear, but preparing flexibly for changing work and responding wisely could also turn it into an opportunity. Since the Industrial Revolution, technology has always been a necessary evil accompanied by a sense of crisis. New productivity has helped overcome the loss of old jobs, and humanity has found ways to transform crises into opportunities.

 

Here are some references I used while writing, along with a few worthwhile reads!

Read next