Can You Build Production Software and AI Without the Hardware?

By Francis Nguyen, Chief Executive Officer··Research
A two-panel house-styled graphic titled "Where the hardware wall actually is." Left: a log-log chart of accelerator memory against model size from one billion to one trillion parameters, with three straight lines for three jobs: training or fully fine-tuning at 16 bytes per parameter, serving at 16-bit at 2 bytes, and serving at 4-bit at half a byte, crossed by dashed reference lines for a 24 GB card, an 80 GB GPU, and an 8-GPU server; at 70 billion parameters the points read 1,120 GB, 140 GB, and 35 GB. Right: a table of the smallest hardware that holds each job for 7B, 70B, and 405B models, in which only three cells have weights that fit on a single card (a 7B model served at 16-bit or 4-bit, and a 70B model served at 4-bit on one 48 GB card, before the KV cache), while training a 70B model needs fourteen 80 GB GPUs across two servers. A note says calling a hosted model API needs no accelerator memory on your side. Exact arithmetic and floors, not a benchmark.

Introduction

Anyone weighing whether to build their own software or AI, rather than buy it, eventually runs into the hardware question. The picture most people carry is racks of GPUs, machines with enormous memory, a data center. Our reading of the research gives a more useful answer than yes or no: it depends on the job, and the job is usually not the one people picture. Hardware is a hard wall in three bounded places. Below them, compute is something you rent by the minute or by the token, and the hardware a growing team is most likely to run into is less glamorous: the cores, memory, and machine time it takes to keep a codebase verified, meaning built and tested automatically on every change, and the discipline to spend them well. A disclosure up front, because it bears on how you read this: Ardenus sells a managed platform, so a conclusion that the hard part of building in-house lies somewhere other than the GPU is seller-convenient. Weigh it accordingly.

  • Hardware is the wall for three jobs: pretraining large models, self-hosting large models at production volume, and the failure and latency behavior that only appears at hyperscale, the size of the largest cloud and AI companies. Those are capital-scale projects, and in early 2026, by one market index’s account, rented GPU capacity was scarce across the board, with half the providers it asked sold out of even the multi-node blocks such projects need.
  • The same model can need 32 times more memory depending on the job. Training takes about 16 bytes of accelerator memory per parameter; serving a model whose weights are compressed to 4 bits takes half a byte. A 70-billion-parameter model that needs fourteen 80 GB GPUs just to hold its training state has 4-bit weights that fit on one 48 GB card.
  • Most applied AI sits below the wall. Hosted models are rented by the token, and the price of a fixed level of benchmark performance has fallen roughly 13-fold a year; the tree models most often used for business prediction train on CPUs; and when RAND asked, nearly all of its industry interviewees said compute was not a limiting factor as long as it was budgeted.
  • The hardware a growing team is most likely to hit is verification. At Google, testing every change individually made compute grow with the rate of change times the size of the test suite; machines starved of cores and memory make tests flaky; and AI-assisted coding raises the volume of change. For most repositories in one large open-source sample that compute was, in our reading, a modest, rentable bill, but it grows with the code and the team, and keeping it trustworthy is an engineering problem as much as a purchase.

Where hardware is the wall

The first wall is pretraining large models from scratch. Epoch AI tentatively estimated in 2024 that the compute used to train notable and frontier models was growing four to five times a year. Cottier and colleagues find that the amortized cost of the final training run for the most compute-intensive models, the share of the hardware’s lifetime cost plus energy that the run used up, has grown 2.4 times a year since 2016, and that buying the hardware outright costs one to two orders of magnitude more: about $800 million for the hardware used to train GPT-4, against $40 million amortized. On trend the largest runs would cost more than a billion dollars by 2027, an extrapolation rather than an observation. Even the famously cheap DeepSeek-V3 run, reported at $5.576 million, is 2.788 million H800 GPU-hours priced at an assumed $2 an hour, and the report itself says the figure covers “only the official training” of the model, “excluding the costs associated with prior research and ablation experiments.”

The second is self-hosting large models at production volume, where the binding resource is memory. A model answering requests keeps a working cache, which engineers call the KV cache, for every token of every request in progress, so memory rather than raw speed caps how many requests one server can handle at once. The vLLM authors calculate 800 KB per token for the 13-billion-parameter OPT model, which uses full multi-head attention, an older and more memory-hungry design, and up to 1.6 GB for a single 2,048-token request. Newer attention designs shrink that figure considerably, but the cache still grows with traffic. At the extreme, DeepSeek describes the recommended deployment unit for its own model as “relatively large, which might pose a burden for small-sized teams”: in its production design, built for online latency targets at high throughput, the minimum unit for generating answers is 40 servers with 320 GPUs, a limitation the report expects more advanced hardware to address.

The third is behavior that only exists at scale. Meta pre-trained Llama 3 405B on up to 16,000 H100 GPUs and, over a 54-day snapshot, recorded 466 job interruptions, 419 of them unexpected, with about 78 percent of the unexpected ones attributed to confirmed or suspected hardware issues. At that size hardware failure is a scheduling fact rather than an incident; Meta reports more than 90 percent effective training time and describes tooling built to raise it. None of these walls is imaginary, but almost no operator building its own software or AI is pretraining a frontier model, serving one to millions of users, or running a sixteen-thousand-GPU cluster.

The memory arithmetic

The chart on this page shows where the wall sits. The memory a model needs is its parameter count times the bytes each parameter costs, and the bytes depend on the job. Training with the standard recipe, mixed-precision Adam, holds 16-bit weights, 16-bit gradients (the corrections computed on each step), and 32-bit optimizer state, which the ZeRO paper tallies as 2 + 2 + 12 = 16 bytes per parameter; for a 1.5-billion-parameter GPT-2 that is “a memory requirement of at least 24 GB,” against the “meager 3 GB” its 16-bit weights need alone. Serving carries no gradients or optimizer state, so the chart counts only the weights for it: 2 bytes per parameter at standard 16-bit precision, or half a byte when each weight is compressed to 4 bits, a technique called quantization, with the KV cache on top. The same model therefore needs 32 times more memory to train than to serve at 4 bits. For a 70-billion-parameter model that is 1,120 GB to train, fourteen 80 GB data-center GPUs across two servers; 140 GB to serve at 16-bit, two GPUs; and 35 GB at 4-bit, one 48 GB workstation card.

The published record matches the arithmetic. QLoRA, which trains small add-on adapters on top of a 4-bit model, cut the memory needed to fine-tune a 65-billion-parameter LLaMA from “more than 780 GB of GPU memory” for regular 16-bit fine-tuning to a single 48 GB GPU. GPTQ quantized the 175-billion-parameter OPT model so that at 3 bits it generates text on one 80 GB A100 where 16-bit needs five, with perplexity, a measure of prediction quality where lower is better, moving only from 8.34 to 8.68. These are floors for the standard recipe, not budgets: activations (the intermediate results a model holds while it computes), the KV cache, and framework overhead all come on top.

The shortcuts have measured limits, which is why the wall is real. The QLoRA authors write that they “did not establish that QLORA can match full 16-bit finetuning performance at 33B and 65B scales,” and that their 33-billion-parameter model “does not quite fit into a 24 GB” card without paged optimizers. A careful comparison on Llama-2-7B found that low-rank adapters, the method QLoRA builds on, learn less than full fine-tuning when a model must absorb a new domain: after further training on math text, the best adapter configuration reached 0.203 on the GSM8K math benchmark against 0.293 for full fine-tuning, though in instruction tuning high-rank adapters could match it. And 3-bit quantization that cost little on the 175-billion-parameter OPT model pushed the 66-billion-parameter model’s perplexity from 9.34 to 14.16. Our reading: teaching a model something genuinely new takes full fine-tuning, whose hardware grows with model size (two 80 GB GPUs at 7 billion parameters, fourteen at 70 billion), as does serving a big model at full precision; adapting or serving a modest one does not.

Most applied AI sits below it

Most of what an operator would build sits well below that boundary. Business prediction on tables, such as churn, demand, or the chance a job runs long, is still most often done with gradient-boosted trees, though tabular foundation models now beat them on some small datasets, a comparison we take up in the regimes of prediction. The point here is narrower: the trees train on CPUs. The LightGBM paper reported per-iteration training up to more than twenty times faster than conventional gradient boosting at almost the same accuracy, on a CPU server with 24 cores and 256 GB of memory, a commodity you can rent by the hour; the paper does not compare against GPU training. For language capability, most teams do not train at all; they call a hosted model. Epoch AI estimated in September 2026 that the cost of reaching a fixed level of performance on its benchmarks has fallen about 47 percent a quarter since 2023, roughly 13-fold a year. That is a price for fixed capability, not a forecast of anyone’s bill: teams that chase the frontier, or run long agentic workloads, can still watch total spend rise.

The people who build and buy these systems point the same way. RAND interviewed 65 people in late 2023, 50 of them in industry and 15 in academia, and from the industry interviews identified five leading root causes of AI project failure: misunderstanding the problem, lacking the data, chasing the latest technology, inadequate infrastructure to manage data and deploy models, and problems too hard for AI. Compute was not one of them. Asked directly, “nearly all of the interviewees stated that compute power was not a limiting factor in their work,” because cloud providers sell it on demand, as long as it is budgeted. RAND names two exceptions: data too sensitive to move to the cloud, and organizations at the edge of AI research training their own large language models, where compute can be rationed internally. The finding predates the 2026 squeeze described below, and RAND’s study excluded projects that only used pretrained language models.

Broader surveys agree. A 2022 CSET survey of mostly academic AI researchers (1.7 percent response rate) found 90 percent rating talent very or extremely important to their most significant project, against 52 percent for large amounts of compute, though 76 percent had revised a project for lack of compute at least sometimes in the previous two years. Among EU firms with 10 or more employees that had considered AI but did not use it, Eurostat’s enterprise survey, fielded in early 2025, found lack of expertise the most-cited reason, at about 71 percent, with cost sixth of eight; the list offered no compute option. Our reading: for these respondents hardware is a friction, and rarely the reason projects fail.

Renting has limits

The advice to just rent it is right as a default, and it is why our essay on the data foundation calls compute a utility. That holds for the workloads most operators run, hosted models and models trained on CPUs, which is where the credit-card parity that essay describes applies; it stops at the three walls above. And a utility bought from a market is still a market. In June 2025 AWS cut on-demand prices for its H100-based instances by 44 percent on Amazon Linux, while noting that the growth in demand for GPU capacity “has outpaced industry-wide supply, making GPUs a scarce resource.” By early 2026 the market had tightened. SemiAnalysis, a commercial research firm whose index is the main public source here, reported its H100 one-year contract price up almost 40 percent, from $1.70 to $2.35 per GPU-hour between October 2025 and March 2026, on-demand capacity sold out across GPU types, and half the providers it asked sold out of even 64-GPU blocks of H100s or H200s. Its account is that sellers held the leverage, preferring multi-year terms and large commitments; our reading is that this weighs most on small buyers who need cluster-scale blocks. Microsoft told investors in July 2026 that demand “continues to exceed available supply,” and TrendForce, another market-research firm, forecast server DRAM contract prices up about 90 percent quarter on quarter in the first quarter of 2026. Discounted capacity comes with conditions: Google Cloud’s Spot VMs run up to 91 percent below on-demand prices for many machine types but can be reclaimed at any time and carry no service-level agreement.

Renting versus owning turns on utilization. Andreessen Horowitz’s widely cited 2021 analysis says that for a new startup or project “the cloud is the obvious choice,” and that at scale its cost “can at least double your infrastructure bill.” 37signals reports cutting its annual cloud bill from $3.2 million to $1.3 million with about $700,000 of owned servers, a self-reported case from a mature company with steady load and an experienced operations crew, while Google’s own authors, in The Data Center as a Computer, conclude from one worked CPU-server example that most companies come out ahead renting once utilization, software, and staffing are counted. None of the three concerns GPUs, so for accelerators the logic carries over but the numbers do not. A small team with bursty needs is the textbook case for renting.

The compute you actually hit

When the GPU is not the constraint, the hardware an in-house build runs into is neither training nor serving. It is verification: the cores, memory, and machine time to build and test a growing codebase on every change, in the automated pipeline engineers call continuous integration, or CI. Even organizations with every advantage report hitting limits here. On an average day described in a 2017 paper, Google’s continuous-testing system ran 800,000 builds and 150 million test runs, and its engineers wrote that testing each commit individually was “not cost effective.” When they had tried, the compute needed grew quadratically: the number of changes submitted and the number of tests to run were each growing roughly linearly, and the compute was their product. Their answer was to batch commits into milestones, cut as often as compute allowed, typically every 45 minutes during peak development time. Google’s engineering book says not every test runs before submission mainly because it is “too expensive,” in engineers’ waiting time as much as machines, and that “even given the amount of compute resources we have,” its test systems “are resource constrained.”

The answer is engineering as much as spend. Google’s build-system chapter notes that distributed builds, once the work is broken into small enough units, let it “complete any build of any size as quickly as we’re willing to pay for.” And much of the spend looks recoverable: in one month of the same Google data, about 91 percent of the test targets affected by changes only ever passed, and a retrospective simulation that skipped tests only distantly connected to the changed code would have saved 42 or 55 percent of test resources, depending on the cutoff. Facebook’s machine-learned test selection halved the infrastructure cost of change-based testing while still reporting over 99.9 percent of faulty changes.

At small scale the same compute is cheap per minute and very uneven. A study of 952 open-source repositories on GitHub Actions found that three quarters averaged only 87 machine-minutes a month, while the heaviest quarter averaged 5,914 and consumed 96 percent of all machine time, almost all of it building and testing. As of October 2026, GitHub’s standard Linux runner for a private repository is 2 virtual CPUs and 8 GB of memory at $0.006 a minute; it rents larger Linux machines up to 96 CPUs and 384 GB; and its only NVIDIA GPU runner is a single Tesla T4 with 16 GB of video memory. Ordinary CPU testing is rentable by the minute at almost any size. Tests that need serious GPUs are poorly served by GitHub’s hosted runners and, in our reading, need self-hosted or separately rented GPU machines.

Faster machines can also make developers more productive, a part of the hardware question that is easy to overlook. In a blind experiment at Google, 15 percent of developers had their builds moved to upgraded build machines, cutting their median build time by about 13 percent, a few seconds on average. They reported small but statistically significant gains, 4 to 5 percentage points, in self-rated productivity and satisfaction, and by the third month were running about one more build and submitting about 24 more lines of code a week.

Starved machines make tests lie

Under-resourced verification is not only slow; it can be untrustworthy. A flaky test is one that passes on one run and fails on the next with no change to the code. Silva and colleagues ran 52 open-source Java, JavaScript, and Python projects 300 times under each of 27 resource configurations and found that 46.5 percent of the flaky tests they observed, 283 of 608, were resource-affected: their failure rate rose when the machine was starved, especially below one CPU core and 1 GiB of memory. Past a modest, project-specific configuration, more resources brought no significant improvement. The finding is scoped to deliberately throttled machines and fairly small suites, and such tests also carry timing assumptions that can be fixed in code, but the implication is direct: a failure on an overloaded machine may be the machine as much as the code.

Flakiness is not cheap at any scale. Google reported in 2016 that about 1.5 percent of its test runs gave a flaky result and almost 16 percent of its tests showed some flakiness; a 2017 follow-up found flakiness correlated with a test’s own size and memory footprint more than with the testing tool, and concluded that the lever is writing smaller tests. One plausible mechanism is easy to create on a single workstation: test runners such as Vitest and Jest default, in single-run mode, to up to one worker per CPU core minus one. Start two or three such runs at once, from several AI coding sessions or several developers sharing one build box, and together they can ask for up to two or three times as many workers as the machine has cores. That this pushes workers into the starved range the research describes is our inference, not a measured result.

More code, more to verify

AI-assisted coding changes the verification load, though this evidence is younger and weaker than the rest. In three field experiments at Microsoft, Accenture, and an anonymous Fortune 100 company covering 4,867 developers, Cui and colleagues estimate that using GitHub Copilot raised weekly pull requests by 26 percent, with a wide standard error of about 10 points; in the two experiments that measured builds, the working paper reports builds up about 38 percent, an estimate driven largely by Accenture’s 316 developers. That was 2022 and 2023 autocomplete, not today’s agents, and builds are a count, not compute-hours. Google’s DORA survey estimated in 2024 that each 25 percent increase in AI adoption was associated with 1.5 percent lower delivery throughput and 7.2 percent lower delivery stability; its 2025 report, using a redefined adoption measure, found throughput positively associated and stability still negative. Both are survey associations, not experiments. DORA’s analysis of 1,110 comments from Google engineers offers a possible mechanism, in human time rather than machine time: the time saved writing code “is frequently re-allocated to auditing and verification.” A study of 806 open-source repositories that adopted the Cursor editor found a burst of output that faded after two months, alongside persistent increases of about 30 percent in static-analysis warnings and 41 percent in code complexity under the authors’ main estimator; other estimators found smaller or no significant increases.

Agents add verification work of their own. A 2026 preprint found coding agents re-running identical tests without changing their patch in roughly half to 83 percent of benchmark tasks, though those repeats averaged under 6 percent of a task’s cost, measured in tokens, not CI machine time. GitHub’s documentation states that coding agents on GitHub “consume GitHub Actions minutes and AI credits,” billable for private repositories on GitHub-hosted runners, and since June 1, 2026, Copilot code reviews on those repositories do too. The productivity side is less settled: in METR’s randomized trial, 16 experienced open-source developers took 19 percent longer on their tasks with early-2025 AI tools, a small study in a demanding setting, and METR calls the signal from its later follow-up, which pointed the other way, unreliable. In a sample of 600 rejected agent pull requests, the most common pattern was a pull request closed with no meaningful human review; failed CI or tests were far less common. Our reading is that the bottleneck for agent-written changes is often people.

No study we found measures how much CI compute an AI-authored change consumes compared with a human one, so the conclusion here is our own inference. If AI raises the volume of change, and verification compute grows with the change rate times the size of the test suite, then AI-assisted teams should expect their verification bill, in machine time and in human review, to grow faster than their headcount. DORA reaches a compatible recommendation from the other direction: investing in robust test automation “may provide a better return on investment than optimizing manual reviews.”

Production is a different machine

Some properties genuinely appear only at scale. In Dean and Barroso’s classic illustration, if each server answers in 10 milliseconds but takes a full second one time in a hundred, a request that must gather answers from 100 such servers takes longer than a second 63 percent of the time. They describe such slow episodes as “unimportant in moderate-size systems” but able to dominate at large scale, and most of their remedies are software, though some of the strongest need extra resources. Most failures in one careful sample, though, did not need a large cluster to reproduce; they needed tests. In 198 user-reported failures from five widely used distributed data systems, Yuan and colleagues found that 98 percent could be triggered on three or fewer nodes, given the right inputs, and 77 percent could be reproduced by a unit test. Of the 48 catastrophic failures, 92 percent came from mishandling non-fatal errors the software had explicitly signaled.

What a lab cannot reproduce is production itself. Netflix’s chaos-engineering team writes that even when a whole system can be reproduced in a test environment, they “still believe in the need to run experiments in production,” because synthetic clients never behave quite like real ones. Google’s ML Test Score makes the same point for models: offline testing cannot by itself guarantee live performance, so new models should be released first to a small, growing share of real traffic, a canary release. Yet in a survey of several dozen Google teams, with Google’s infrastructure at hand, none of the rubric’s tests was implemented by more than 80 percent of teams. Our reading: access to that infrastructure did not, by itself, produce the testing discipline.

What this means for the operator

So, can you build production software and AI without the hardware? It depends on which of three kinds of work your plan involves. If it involves pretraining anything beyond a small model, fully fine-tuning a large one, or serving a large open model to many users at full precision, hardware is not an input to the project; it is the project, priced like capital and, in SemiAnalysis’s early-2026 snapshot, rented on terms sellers set. If it is applied AI, meaning hosted language models plus prediction models trained on your own operational data, you may not need to own or run GPUs at all, which is the note beneath the chart. The scarce inputs the surveys point to are expertise and data; the data side is the subject of the data foundation beneath enterprise AI.

Either way, a team building production software runs into the third kind of work: verification. Budget it as a real cost that grows with the code and the team. Give test runs real machines, well clear of the under-one-core, under-1-GiB range where resource-affected flakiness concentrates. Spend on test selection, caching, and build tooling before brute force. And keep canary releases and monitoring, because no test bench reproduces production. A team that does those things can build serious software on rented machines; a team that skips them will not be rescued by a bigger GPU. If the tools in question are AI app builders rather than a codebase of your own, their separate limits are covered in the governance ceiling of vibecoded apps.

The disclosure from the top belongs here in full. Ardenus sells a managed platform, so an argument that the hard part of building in-house is verification and production engineering, rather than buying GPUs, is seller-convenient, and a reader should weigh it that way. To be fair to the other side: renting compute works and is the right default, the price of a fixed level of AI capability is collapsing, and the hard part we point to is tractable too. For three quarters of the open-source repositories in the largest CI study here, verification compute was, in our reading, a small bill; test selection recovers much of the rest, as Facebook’s halving of its change-based testing cost shows; and METR describes its later, more favorable estimate of AI’s effect as likely a lower bound. A disciplined small team can build production software without owning hardware. Two limits on the evidence matter as well: almost none of it measures small commercial organizations directly, so applying it to an operator is an extrapolation, and the market figures here are dated and will move. No result, saving, or metric is attributed to Ardenus here, no client or first-party data appears anywhere in this essay, and the chart is exact arithmetic about memory rather than a measurement of anyone’s systems. You can read more of our research on the Ardenus articles hub, or see the platform itself on the technology page.

Sources and methodology

This essay was researched with a multi-agent sweep of six angles and two targeted follow-ups, which gathered roughly 150 candidate claims from primary sources, followed by an adversarial fact-check that re-downloaded each load-bearing source and confirmed every quoted passage against its raw text. Of the 82 claims checked, none was refuted and 49 were narrowed or corrected, most often to restore a hedge, a sample, a version, or a date; the essay uses each at its corrected scope, and a final review mapped every number and quotation in the draft back to that verified record. The chart is exact arithmetic with no external data beyond the bytes each parameter costs: 16 for mixed-precision Adam training (ZeRO), 2 for 16-bit weights, and half a byte for 4-bit weights, multiplied by the parameter count in decimal gigabytes. Its figures are floors for that standard recipe that exclude activations, the KV cache, and framework overhead, and the committed build script node:assert-locks them, including three published figures the formula reproduces exactly (ZeRO’s 24 GB for a 1.5-billion-parameter GPT-2, and GPTQ’s five 80 GB GPUs at 16-bit and one at 3-bit for OPT-175B) and one it is consistent with (QLoRA’s 65-billion-parameter fine-tune on one 48 GB GPU, where the 4-bit weights alone take 32.5 GB). Market figures (GPU rental prices, memory prices, runner specifications and prices) are dated to their sources, and single-source commercial data from SemiAnalysis and TrendForce is identified as such. Statements marked as our reading or our inference are the essay’s own, including the link from AI-assisted coding to verification compute, which no study we found measures directly. No client or first-party data is used, and no result, saving, or forecast is attributed to Ardenus.

  1. Training compute of frontier AI models grows by 4-5x per year (Sevilla and Roldán, Epoch AI, 2024) - a tentative 4-5x a year; 4.1x a year for notable models, 2010 to May 2024.
  2. The Rising Costs of Training Frontier AI Models (Cottier et al., arXiv 2405.21015, preprint) - amortized final-run cost growing 2.4x a year since 2016; about $800M to acquire GPT-4’s training hardware against $40M amortized; the $1B by 2027 figure is an extrapolation.
  3. DeepSeek-V3 Technical Report (DeepSeek-AI, arXiv 2412.19437) - 2.788M H800 GPU-hours at an assumed $2 an hour for the official run only, and a recommended deployment unit that might burden small teams.
  4. Efficient Memory Management for Large Language Model Serving with PagedAttention (Kwon et al., SOSP 2023) - 800 KB of KV cache per token calculated for OPT-13B, up to 1.6 GB per 2,048-token request.
  5. The Llama 3 Herd of Models (Llama Team, AI at Meta, arXiv 2407.21783) - 466 interruptions over a 54-day snapshot on up to 16K H100 GPUs, about 78% of the 419 unexpected ones attributed to confirmed or suspected hardware issues; effective training time above 90%.
  6. ZeRO: Memory Optimizations Toward Training Trillion Parameter Models (Rajbhandari et al., SC20) - 16 bytes of model state per parameter for mixed-precision Adam; at least 24 GB for GPT-2 1.5B.
  7. QLoRA: Efficient Finetuning of Quantized LLMs (Dettmers et al., NeurIPS 2023) - a 65B model fine-tuned on a single 48 GB GPU against more than 780 GB for 16-bit fine-tuning; parity at 33B and 65B not established.
  8. GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers (Frantar et al., ICLR 2023) - 3-bit OPT-175B on one 80 GB GPU against five at FP16.
  9. LoRA Learns Less and Forgets Less (Biderman et al., TMLR 2024) - adapters trail full fine-tuning in continued pretraining on code and math.
  10. LightGBM: A Highly Efficient Gradient Boosting Decision Tree (Ke et al., NeurIPS 2017) - up to more than 20x faster per training iteration on a 24-core, 256 GB CPU server.
  11. The plunging price of thought (Emberson and Roodman, Epoch AI, September 2026) - the cost of a fixed level of benchmark performance falling about 47% a quarter since 2023.
  12. The Root Causes of Failure for Artificial Intelligence Projects and How They Can Succeed (Ryseff, De Bruhl and Newberry, RAND, 2024) - five root causes, and compute not a limiting factor for nearly all industry interviewees when budgeted.
  13. “The Main Resource is the Human”: A Survey of AI Researchers on the Importance of Compute (Musser et al., CSET, 2023) - fielded in 2022; 410 complete responses, 1.7% response rate, mostly academic.
  14. Use of artificial intelligence in enterprises (Eurostat, Statistics Explained, survey fielded in early 2025) - reasons for not using AI among EU firms with 10 or more employees that considered it.
  15. Announcing up to 45% price reduction for Amazon EC2 NVIDIA GPU-accelerated instances (AWS News Blog, June 2025) - the 44% on-demand cut; 45% is a Savings Plan rate.
  16. The Great GPU Shortage: Rental Capacity (SemiAnalysis, April 2026) - a commercial rental-price index; a single source.
  17. Fiscal Year 2026 Fourth Quarter Earnings Call (Microsoft Investor Relations, July 29, 2026).
  18. TrendForce press release (February 2, 2026) - a memory contract price forecast from a market-research firm.
  19. Spot VMs (Google Cloud Compute Engine documentation, as of October 2026).
  20. The Cost of Cloud, a Trillion Dollar Paradox (Wang and Casado, Andreessen Horowitz, 2021).
  21. Our cloud-exit savings will now top ten million over five years (Heinemeier Hansson, 37signals, 2024) - a self-reported case.
  22. The Data Center as a Computer, Fourth Edition (Barroso, Hölzle and Ranganathan, Springer, published December 2025, © 2026) - a worked rent-versus-own example for one CPU server.
  23. Taming Google-Scale Continuous Testing (Memon et al., ICSE-SEIP 2017).
  24. Software Engineering at Google (Winters, Manshreck and Wright, O’Reilly, 2020), chapter 23 on continuous integration and chapter 18 on build systems.
  25. Predictive Test Selection (Machalica et al., ICSE-SEIP 2019).
  26. Resource Usage and Optimization Opportunities in Workflows of GitHub Actions (Bouzenia and Pradel, ICSE 2024).
  27. GitHub Docs, as of October 2026: GitHub-hosted runners, larger runners, and Actions runner pricing.
  28. Developer Productivity for Humans, Part 4: Build Latency, Predictability, and Developer Productivity (Jaspan and Green, IEEE Software, 2023).
  29. The Effects of Computational Resources on Flaky Tests (Silva et al., IEEE Transactions on Software Engineering, 2024).
  30. Google Testing Blog: Flaky Tests at Google (Micco, 2016) and Where do our flaky tests come from? (Listfield, 2017).
  31. Test-runner worker defaults: Vitest (source, v5.0.3) and Jest (source, v30.5.2).
  32. The Effects of Generative AI on High-Skilled Work: Evidence from Three Field Experiments with Software Developers (Cui et al., Management Science, 2026; build figures from the working paper).
  33. DORA (Google Cloud): Accelerate State of DevOps Report 2024, the 2025 report announcement, and Balancing AI tensions (2026).
  34. Speed at the Cost of Quality: How Cursor AI Increases Short-Term Velocity and Long-Term Complexity (He et al., MSR 2026).
  35. Analyzing and Mitigating Cost-Inefficient Behaviors in Coding Agents (Hu et al., arXiv 2609.30725, preprint).
  36. GitHub: About third-party coding agents and Copilot code review consuming Actions minutes from June 1, 2026.
  37. Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity (Becker et al., METR, arXiv 2507.09089, preprint, 2025) and the February 2026 update.
  38. Where Do AI Coding Agents Fail? An Empirical Study of Failed Agentic Pull Requests in GitHub (Ehsani et al., MSR 2026).
  39. The Tail at Scale (Dean and Barroso, Communications of the ACM, 2013).
  40. Simple Testing Can Prevent Most Critical Failures (Yuan et al., OSDI 2014).
  41. Chaos Engineering (Basiri et al., IEEE Software, 2016).
  42. The ML Test Score: A Rubric for ML Production Readiness and Technical Debt Reduction (Breck et al., IEEE Big Data 2017).