Government-Industry Distance and Japan's Late Start in Generative AI: The 2001 ADSL War, and Why Having TSUBAME Didn't Produce a Large Language Model
By 2008, Japan was running one of the world's first GPU supercomputers. A large language model still never came out of it. Was that a hardware problem, or an organizational one?
Framing the Question
"Japan fell behind on generative AI" has been repeated so often since 2023 that it has stopped being an analysis and become a stock phrase. And as stock phrases tend to do, it drags sloppy explanations along with it: there wasn't enough compute, there weren't enough people, there was no way to match the volume of English-language data.
This article starts from the observation that one of those explanations — "there wasn't enough compute" — is factually rather shaky. When it comes to using GPUs for scientific computing, Japan was well ahead of the curve. And yet that compute was never turned toward training a large foundation model. If that's the case, then the problem wasn't whether the hardware existed, but rather the machinery that decides how the hardware gets used. That is this article's thesis.
Before getting to it, though, we'll set up a control case: an episode in which an outsider actually did break an entrenched monopoly structure in the Japanese tech market. That's the ADSL war of 2001.
The Control Case: When a Fixed-Line Monopoly Actually Broke, in 2001
1999–2000: Regulation Moved First
In the late 1990s, internet access in Japan depended on NTT's subscriber telephone lines — copper. The rules that would let other carriers use that "last mile" — unbundling of the local loop — and the rules that would let them install their own equipment inside NTT's exchange buildings — collocation — were developed by what was then the Ministry of Posts and Telecommunications and later the Ministry of Internal Affairs and Communications. Between 1999 and 2000, new entrants such as Tokyo Metallic Communications began offering ADSL service over leased NTT lines.
Having rules on paper and having installation work actually proceed, however, turned out to be two different things. It was widely observed at the time that NTT spent a very long period "testing" new entrants' DSL equipment, and that the paperwork for in-exchange collocation was extraordinarily cumbersome. The level of interconnection fees and the scope of unbundling also became a point of contention in the 2000 Japan-U.S. deregulation talks; in practice, it took several forces at once — foreign pressure included — before the terms finally moved.
2001: Price Destruction and the Red Bag
That's the market SoftBank entered. On June 19, 2001, Yahoo! JAPAN announced "Yahoo! BB," a broadband service headlined by an ADSL connection fee starting at 990 yen per month, and launched the service that September. At 2,280 yen per month (pre-tax) including ISP service, for up to 8 Mbps downstream, the pricing was unambiguous price destruction by the standards of the day. The going rate then was 5,000 to 6,000 yen a month for an ADSL line plus provider fees; after Yahoo! BB arrived and competitors followed, that was pushed down into the 3,000–4,000 yen range. By December of the same year, NTT itself had cut FLET'S ADSL pricing into the 2,000-yen range.
Paired with that pricing strategy, and just as memorable, was the free distribution of ADSL modems on the street. Booths under red-and-white parasols were staffed by what came to be called the "parasol squad," handing modems in red paper bags to passers-by — guerrilla promotion of a kind highly unusual for a telecom carrier, rolled out in front of train stations nationwide. As a footnote: that Yahoo! BB ADSL service ran for 22 years and was finally shut down at the end of March 2024.
What Was Happening Inside the Exchange Buildings
Signing up subscribers means nothing, though, if the interconnection work inside NTT's exchange buildings doesn't get done. This is where SoftBank ran into the slow collocation process described above. The anecdote of Masayoshi Son marching into the ministry himself and confronting officials in extremely forceful terms has been retold in numerous business books and media accounts, typically with the conclusion that NTT's installation work started moving shortly afterward.
The structure that can be extracted from the 2001 case is this. (1) The regulator first created rules opening up the incumbent's monopoly assets (the copper and the exchange buildings). (2) An entrant from outside the existing telecom industry then came in with pricing and promotional tactics no established player would ever adopt. (3) When the entry barrier turned out to still be operating at the practical level, someone pushed back against it directly with the regulator. The result moved prices across the whole market, incumbent included. Genuine disruptive competition can happen in Japan's tech market when the conditions line up.
The Main Question: Did Japan Have the Compute?
Now to generative AI. First, let's check with facts how far the "there wasn't enough compute" explanation actually holds.
TSUBAME: A GPU Supercomputer Ahead of the World
The TSUBAME series at the Tokyo Institute of Technology (now Institute of Science Tokyo) was one of the world's earliest supercomputers to seriously commit to using GPUs for scientific computing.
- TSUBAME1.0 (operational April 2006): peak performance on the order of 85 TFLOPS; the fastest machine in Japan and Asia at the time
- TSUBAME1.2 (October 2008): 170 NVIDIA Tesla S1070 units added (680 GPUs). Described as the first GPU-equipped supercomputer to rank among the upper tier of the TOP500, taking peak performance from roughly 80 TFLOPS to roughly 141 TFLOPS
- TSUBAME2.0 (2010): Japan's first petaflops-class machine, and ranked No. 2 worldwide on the Green500 power-efficiency list
- TSUBAME3.0 (2017): 540 nodes with 2,160 NVIDIA Tesla P100 GPUs; ranked No. 1 worldwide on the Green500
- TSUBAME4.0 (operational April 1, 2024): 240 nodes, each with two AMD EPYC processors and four NVIDIA H100 GPUs (960 GPUs in total)
The year that matters here is 2008. Not long after NVIDIA's CUDA appeared, a Japanese university was running hundreds of GPUs together in production. Deep learning wouldn't draw broad attention until the image-recognition breakthrough of 2012 and after — which means that in terms of GPU-computing infrastructure and operational know-how, Japan was on the leading side.
Fugaku and ABCI: The Public Money Was Spent
National-level investment was not small either. Fugaku, developed by RIKEN and Fujitsu, took the No. 1 spot on the TOP500 in June 2020 and entered full operation in 2021. The government budget covering R&D, procurement, and application development has been put at roughly 110 billion yen. But Fugaku's architecture is built on the Fujitsu A64FX (Armv8.2-A with SVE) in a CPU-only configuration, with no GPUs. That was a deliberate design choice in the HPC context of its time, but it left the machine ill-suited to large-scale Transformer training in any straightforward way. It is not unrelated that "Fugaku-LLM," released in May 2024, came in at 13 billion parameters. (The fact that the effort began with porting Megatron-DeepSpeed so that Transformers would run on Fugaku at all tells the same story.)
As for compute built specifically for AI, the National Institute of Advanced Industrial Science and Technology brought ABCI (AI Bridging Cloud Infrastructure) into full operation in August 2018, with the first generation built on NVIDIA Tesla V100. Since January 2025, ABCI 3.0 has been running with 6,128 NVIDIA H200 GPUs, serving as a national AI computing base open to universities, public research institutes, and startups. The Ministry of Economy, Trade and Industry and NEDO also launched "GENIAC" in February 2024, which has continued to provide compute support to foundation-model development projects (the first round selected 10 projects, with subsequent rounds selecting somewhere in the range of 16 to 24 each).
So What Was Actually Missing?
Line up the numbers and the situation comes into focus. GPT-3 was announced in June 2020 at 175 billion parameters; GPT-4 in March 2023. Over that period, U.S. frontier labs were committing thousands to tens of thousands of GPUs to a single training run. TSUBAME4.0, which went live in 2024, has 960 GPUs — and it is a shared resource supporting research and education across an entire university. ABCI 3.0's 6,128 GPUs are likewise a national shared platform divided among users nationwide.
In other words, "Japan didn't have a single GPU" is plainly false — but there was effectively no resource that could be used the way frontier labs used theirs: one team monopolizing thousands of GPUs for months to pour into a single speculative training run. And that gap is less a difference in total hardware than a difference in how the resource gets allocated. From here on, the story is about organizations.
Structural Hypotheses: What Was on the Organizational Side
(1) Shared-Allocation Models Fit Speculative Large-Scale Training Badly
Many of Japan's national supercomputers are operated under frameworks such as HPCI (the High Performance Computing Infrastructure), in which research proposals are solicited publicly, a review committee selects among them, and compute is allocated to the accepted proposals. As a mechanism for distributing publicly funded resources across disciplines, without favoring particular actors, on defensible criteria, this is entirely reasonable.
But the mechanism is optimized for computation whose results can be justified in advance. A proposal must state objectives, methods, and expected outcomes; it goes through review; and the compute time awarded comes with caps and a fixed period. That does not mesh — starting from the very premises of the review — with a scaling-law experiment along the lines of "we honestly don't know what happens if we scale this architecture up this far, but we'd like to occupy a cluster for several months and find out." And training a large language model is a process with high failure costs: if it diverges partway through, you start over.
(2) Risk-Averse Public Funding vs. the VC-Style Single Bet
The U.S. model that produced foundation models from GPT-3 onward was, put bluntly, this: a company with no revenue yet burns a large sum raised from venture capital on one large-scale training run of uncertain probability of success. For that to work, you need a layer of capital deep enough that failure still leaves room for the next bet.
Japan's funding environment differs in exactly that depth. In the first half of 2024, VC investment was about 14.2157 trillion yen in the United States against about 144.7 billion yen in Japan. Japanese startups raised roughly 761.3 billion yen across 2024 — reported at around 0.11% of GDP. The gap is partly one of absolute size, but the more fundamental point is that the practice of putting tens of billions of yen into one company before it has a product is hard to sustain, institutionally and culturally. That goes double for public money: when the source is tax revenue, outcomes must be justified up front and failures must be accounted for afterward. This is not a defect in the system; it is the system behaving as designed. It simply fits speculative large-scale training badly.
(3) The Capacity to Procure and the Capacity to Keep Running Are Different Things
Another point that tends to get overlooked: installing a supercomputer and finishing a foundation model demand completely different organizational capabilities.
Procuring a large machine is a self-contained sequence — securing budget, drafting specifications, running a tender, installing, benchmarking — and the outcome is made visible as a ranking on the TOP500 or Green500. This is a domain in which Japanese government and academia have been highly practiced for decades; TSUBAME and Fugaku are the proof. But it carries a structural side effect: once the ranking is published, the project tends to be treated as complete.
Developing a foundation model, by contrast, has to keep running for years — data collection and cleaning, failed runs and restarts, evaluation, safety work, building serving infrastructure, and shipping it as a product — with no clear moment of "completion" anywhere along the way. What it requires is not the ability to acquire hardware but an organization that can carry an uncertain, long-horizon project, failures included. If Japan was strong at the former while actors capable of the latter struggled to emerge, that is not a problem about computers.
(4) Concentration of Public Funding as a Structural Pattern
This next point needs to be handled more carefully. A structural tendency is frequently noted about how public money flows into AI and advanced-technology work in Japan: that grants and public-private partnership contracts tend to concentrate toward companies spun out of a small number of prominent national-university laboratories. This article treats that strictly as a general observation about structure, without pointing to any specific company, laboratory, or individual.
There are rational reasons for such concentration. The number of organizations with the track record and institutional capacity to be trusted with large-scale technical development is limited to begin with, and they are easier for reviewers to evaluate. It is also a perfectly defensible judgment that concentrating a limited budget on actors with proven execution ability produces results more reliably than spreading it thinly.
On the other hand, it has been argued that concentration can carry costs in competitive dynamics. Fewer independent bets means less diversity among the technical approaches that get explored. When the set of eligible recipients is effectively narrow, "going after public funding" becomes, for a new entrant, less a matter of validation in the market and more a matter of connecting into an existing network. And to the extent that funding decisions become influenced by academic pedigree and personal networks rather than market track record, their precision as a selection signal declines.
To be explicit: this is not an assertion that any misconduct or improper relationship occurred. This article identifies no specific problem with any individual grant or contract award, and is not in possession of any such information. What is being described is a general point about institutional design — that a system in which resource allocation structurally concentrates can, independent of anyone's intentions or good faith, push in the direction of reducing competitive diversity.
Placing 2001 and the 2020s Side by Side
Return to the control case. Disruptive competition succeeded in the 2001 telecom market because three things came together.
- The monopoly asset had been opened up institutionally: rules for local-loop unbundling and collocation were in place first, if imperfectly
- Capital from outside the industry entered while ignoring the industry's conventions: an executive who had not come up through telecom brought in pricing and promotional tactics no incumbent would ever have chosen
- Someone pushed back against the barriers at the practical level: the party affected took the failure of the rules directly to the regulator
Now apply those three conditions to generative AI development in the 2020s. In this author's reading, none of them held easily. The key resource — large-scale compute that one party can occupy exclusively — was never held by a monopolist in the first place; it existed as public shared infrastructure, distributed "fairly." Since it does not take the shape of a monopoly structure to be opened up, the 2001 move of institutional opening has nothing to grip. The layer of capital willing to bring in vast sums from outside the industry is thin. And there was no clear wall of vested interest to push back against either.
Put differently: where the problem in 2001 was that the wall was too high, the problem in the 2020s may have been that there was no wall, but also no mechanism for making big bets. The first kind of problem can be broken with regulation. The second cannot.
If that reading is right, the implication is not especially pessimistic — because it does not say that Japan's tech market is structurally incapable of absorbing disruptive competition. 2001 shows that it happens when the conditions align. What is needed is not only procuring more machines, but a design question: how do you build actors that can allocate exclusive resources and time to long-horizon projects that will probably fail?
Summary
- Yahoo! BB's 2001 entry is a case where institutional groundwork (local-loop unbundling) and price-destroying entry from outside the industry combined to actually move NTT's fixed-line monopoly
- With TSUBAME1.2 in 2008, Japan fielded what is described as the first GPU-equipped supercomputer in the upper tier of the TOP500 — it was on the leading side of GPU computing
- Fugaku (2020, roughly 110 billion yen in public funds) is a CPU-only design, not one that lends itself straightforwardly to large language model training
- Public AI compute and support programs do exist: ABCI (from 2018; version 3.0 went live in 2025 with 6,128 H200 GPUs) and GENIAC (from 2024)
- But all of these are shared, allocated resources, ill-suited to the pattern of one team occupying a cluster for months to run a speculative large-scale training job
- The primary cause of the late start is therefore likely not the total amount of compute, but organizational structure: how resources are allocated, the risk tolerance of public funding, the gap between procurement capability and long-term operational capability, and the thinness of the independent VC layer
- The structural tendency for public funding to concentrate on spinouts from a small number of prominent university labs is sometimes discussed as a factor reducing competitive diversity, but this is one interpretation and is not an allegation of any specific wrongdoing
- The question as a whole is multi-causal, and this article's structural hypotheses should be weighed alongside the other candidate explanations
On Technical Networks and Informational Edge
A structure in which funding and resource allocation is shaped by personal networks can also be analyzed through the lens of information asymmetry.
Read the Techno-Insider Theory