Is infra ready for recursive self-improvement?
TL;DR. An RSI loop runs at the speed of its verifier, not its proposer.
- Verification cost grows by orders of magnitude as the target moves from a kernel to a whole model to a training run. Once the target involves training, verification fills almost every iteration, and faster generation barely moves the loop.
- Latency can be bought; noise cannot. The runs needed for a reliable verdict grow like \(\sigma^2/\delta^2\), and gains shrink near the frontier.
- Building on verifier’s rule: verifiability is not fixed by the task. It is partly a property of infrastructure, so infrastructure can make RSI’s task more verifiable: a calibrated fidelity ladder, content-addressed caching, weight inheritance and boundary graph replay, paired controls, and a hardened verifier.
- Beyond code, the limit is measurement. RSI in the physical world is bounded by the precision and throughput of its instruments.
Every recursive self-improvement (RSI) loop has the same shape: propose a change, verify it, keep what works, and (in the learning variants) update the proposer on the outcome. Most of the discussion focuses on the proposer: bigger models, better prompts, RL on the operator. My argument in this post is that the proposer is not where the loop will stall. It will stall at verification, and the infrastructure we have today was built for a world where verification is cheap.
Starting point: verifier’s rule
Jason Wei’s asymmetry of verification and verifier’s rule is the right place to start. Some tasks are much easier to verify than to solve, and his rule is that “the ease of training AI to solve a task is proportional to how verifiable the task is.” He names five properties that make a task verifiable: objective truth, fast to verify, scalable to verify, low noise, and continuous reward. Guess-and-check systems like AlphaEvolve work best when all five hold.
I take the rule as given and push it one step further. In RSI, the task is improving the machine-learning system itself, and that is where all five properties are weakest. But verifiability is not only a property of the task. It also depends on the infrastructure that runs the checks, and each property has an infrastructure lever:
| Property | Where RSI tasks fall short | Infrastructure lever (below) |
|---|---|---|
| Objective truth | the validation score is a proxy, and the verifier can be gamed | hidden evaluation, adversarial hardening, calibration against ground truth |
| Fast to verify | a training run takes hours to weeks | fidelity ladder, weight inheritance, boundary graph replay, caching |
| Scalable to verify | training is GPU-bound, sandboxes are CPU-bound | content addressing, heterogeneous scheduling, asynchronous updates |
| Low noise | seed noise, the gain from simply training longer, and in the physical world the instrument’s noise floor | null-edit siblings, interleaved A/B timing, and better instruments |
| Continuous reward | near the optimum, gains are too small to rank reliably | paired comparisons, rank calibration with Kendall \(\tau\) |
Verifier’s rule predicts which tasks AI will solve first. The infrastructure question is how far we can move the task RSI cares about up that list.
A simple accounting
Take a parent solution \(x\) with fitness \(f(x)\) and a proposer \(q\) that samples children \(y\). Let \(p(x)\) be the probability that a child beats its parent, and \(\mathbb{E}[\Delta \mid \text{success}]\) the expected gain when it does. The number of verifications needed to climb from \(f_0\) to \(f^*\) is roughly
\[V \approx \int_{f_0}^{f^*} \frac{df}{p(f)\,\mathbb{E}[\Delta \mid f]}.\]Everything people do to the proposer (in-context archives as in FunSearch and AlphaEvolve, RL on the weights, larger models) is an attempt to raise \(p\). RL now shows up in two forms: test-time RL on a single problem, as in ThetaEvolve and TTT-Discover, and post-training the search operators themselves, as Frontis-MA1 does in the open OpenRSI stack, where the Draft, Improve, Debug, and Crossover operators are trained with execution-grounded SFT and online RL and then composed into long-horizon search. Everything people do to the evaluator (cascades, early stopping, cheaper proxies) is an attempt to lower the cost of each term in \(V\).
There is a floor. A perfect proposer would jump straight to the optimum and need a single verification. Real proposers cannot, because some of the uncertainty about how a training run behaves on a specific dataset cannot be removed by a better prior. You only find out by running it. That residual uncertainty is what makes verification a hard lower bound, and it suggests a one-line summary:
The asymptotic rate of verified self-improvement is verification throughput times the useful information per verification.
Wall-clock has its own floor: the serial depth of the improvement chain times the latency of one verification.
The loop runs at the speed of verification
What one verification costs depends on what is being verified. Three concrete targets, from cheap to expensive:
- A single kernel. For example, an attention forward kernel such as FlashAttention-4 on a B200. One call takes milliseconds. Verifying a candidate means compiling it, checking its numerics against a reference over many shapes, and timing it repeatedly, so the check costs far more than the call.
- A whole-model kernel. For example, the Llama-3.2-1B megakernel from Hazy Research, which fuses the entire forward pass into one kernel. Verifying a change means decoding an eval set of prompts, checking the outputs against a reference model, and measuring end-to-end tokens per second.
- A training change. Here the verifier is a training run. Karpathy’s llm.c reproduction trains GPT-2 124M in about 90 minutes on 8 A100s. LLaMA-65B took about 21 days on 2,048 A100s.
Public timing reference points
I do not have measurements of a complete verification loop for each domain. The table below instead collects timings from papers and first-party reports. Each row has a specific scope: a kernel forward pass, full-model decode per output token, generation or search for one problem, an agent’s complete issue attempt, or a training run. These are useful anchors for a cost model, but their different workloads and hardware do not form a controlled comparison of verification latency.
| Workload | Time shown | Configuration and timed scope | Primary source |
|---|---|---|---|
| FlashAttention-4 | ≈5.45 ms | 1×B200; BF16, non-causal forward, batch 1, sequence 32,768, 16 heads, head dimension 128. Converted from 1,613 TFLOP/s; excludes compilation and correctness testing. | Paper §5, Fig. 4 |
| DeepSeek-R1 decode | ≈9.01 ms/token | 8×B200; one concurrent request, 1K input / 2K output tokens, BF16 attention and NVFP4 FFNs. The no-MTP configuration reports 111 tokens/s/user; this is amortized time per output token, not a complete request or eval set. | TensorRT-LLM report, MTP ablation |
| MATH500 search | 11.2 s/problem, mean | Qwen2.5-7B-Instruct, beam width 4, with a process reward model; experiments on a 3×A100 node, one model per GPU. Search latency, not an isolated answer or proof check. | SPECS, Table 1 and Appendix E.2 |
| HumanEval generation | 50.94 s/example, mean | Autoregressive baseline in Discrete Flow Matching, batch 1 on 1×A100 80GB. Generation latency; unit-test execution is not the reported metric. | Paper, Appendix F |
| SWE-bench Lite attempt | 309 s/issue, mean | SpecRover, Claude 3.5 Sonnet with GPT-4o fallback; 300 issues. Includes generation, reproducer tests and regression tests; no isolated test-runtime or hardware breakdown. | SpecRover, Table I and RQ2 |
| Train GPT-2 124M / ≈10B tokens | ≈90 min | 8×A100 80GB SXM; llm.c reproduction trained from scratch on FineWeb. Author-reported run duration. | Karpathy’s reproduction report |
| Train Goedel-mHC 1B / 20B tokens | ≈21 h | 8×H200 SXM; BF16, FineWeb-Edu, sequence length 4,096. Author-reported training wall time. | Model card, Training |
| Train Olmo Hybrid 7B / 3T tokens | 7 days | 512×B200 across 64 HGX systems. Reported training duration. | Lambda / Olmo training report |
| Train LLaMA 65B / 1.4T tokens | ≈21 days | 2,048×A100 80GB; the paper gives this approximate duration from its training throughput. | LLaMA, §2.4 |
The two B200 conversions are explicit. For FlashAttention, the reported forward-pass work gives \(t = \frac{4\cdot 32768^2\cdot 128\cdot 16}{1613\cdot 10^{12}} \approx 5.45\,\mathrm{ms}\). For DeepSeek, \(t = 1/111\,\mathrm{s} \approx 9.01\,\mathrm{ms}\) per output token. At that rate, 2K output tokens take about 18 seconds; prompt processing and a full correctness suite require separate accounting. I use the no-MTP ablation so the unit is an ordinary decode token; the report’s faster speculative-decoding configurations are different workloads.
A millisecond kernel pass is only one part of a candidate evaluation: compilation, reference comparisons, input coverage and repeated timing still have to be paid for. Similarly, evaluating a training change includes the training run and the held-out evaluation. The sources above do not establish a verification fraction, a p90, or the number of verified improvements possible per day. Those require measurements of the actual loop. The training examples also change GPU count, architecture and token budget, so their durations are examples rather than a scaling law.
For scheduling, the distinction matters: a synchronous batch waits for its slowest member, and a SWE-style attempt uses both model inference and a test environment. Profiling those components separately is what tells us whether to invest in faster inference, CPU sandboxes, or training capacity.
What that does to the loop
From the first target to the last, the proposal stays about the same size, a diff to some code, while the time to verify it grows by many orders of magnitude. The diagram below shows what that does to the loop. Each row is the same wall-clock window filled with iterations of propose (amber) then verify (teal). Pick a task, then speed up one side.
The diagram is Amdahl’s law applied to the loop. Speeding up the side that dominates an iteration is the only speedup that counts, and for anything that involves training, that side is verification. A cheap proxy helps, but only up to the rounds that still need ground truth.
Three costs grow together as the target moves toward the learner itself:
- Latency. A full training run can keep a candidate waiting for hours or days. Faster proposal generation cannot remove that dependency.
- Information per verification. A proof check is a clean bit. An ML score carries seed noise, and near the optimum the improvements \(\delta\) you are trying to detect shrink. The next section is about this cost.
- Proxy gap. In math the verifier is the objective. In ML the validation score is a proxy for test performance, and a small-scale run is a proxy for a large-scale one. Proxies need periodic calibration against ground truth.
The diagram shows only the first cost. The other two multiply it, and the second one deserves its own section.
Noise: the limit that compute does not buy
Latency can be bought: more GPUs, more sandboxes, caching, replay. Noise is different. A verification is a measurement, and the information it carries is set by the ratio between the effect you are looking for and the noise around it. When a child’s true gain is \(\delta\) and one run has noise \(\sigma\), the number of runs needed for a reliable verdict grows like \(\sigma^2/\delta^2\). Near the frontier the gains get smaller, so every halving of \(\delta\) costs four times as many runs.
A paired control means the child and its reference are measured under the same randomness: the same seed, the same data order, the same hardware, run back to back. For training, the reference is the null-edit sibling described below; for kernels, it is the parent timed in interleaved A/B runs on the same GPU. Noise that both runs share then cancels when you subtract them. If a fraction \(\rho\) of each run’s variance is shared, the variance of the measured difference falls from \(2\sigma^2\) to \(2(1-\rho)\sigma^2\). The diagram uses \(\rho = 0.9\), and a “reliable verdict” means a one-sided test at 5% false positives with 80% power. The exact constants do not matter. The slope does: no amount of parallel hardware changes the \(1/\delta^2\) law, it only pays for it.
In a digital loop much of the noise is ours to control. We can fix seeds, reuse the same data order for parent and child, run a null-edit sibling, and interleave A/B timings on one GPU. That is why the paired curve exists at all, and why the infrastructure sections below spend so much effort on it.
The physical world does not offer those controls. The noise floor is set by the instrument, and a run cannot be replayed with the same random numbers. Particle physics shows the extreme. The LHC collides proton bunches 40 million times per second, and the ATLAS trigger keeps a few hundred of those events per second; the same note puts the Higgs production cross-section about ten billion times below the total. Claiming the Higgs took ATLAS and CMS a five-sigma signal built from that stream. Radio astronomy is similar: the 2017 Event Horizon Telescope campaign recorded about 4 PB to produce a single image of the M87 black hole. In both cases the bottleneck was never the volume of data. It was how much signal survives the noise.
This is the reason RSI does not extend easily from code into the physical sciences. A loop that proposes materials, drugs, or experiments is verified by measurements, and its rate of verified improvement is bounded by the precision and throughput of its instruments. If the infrastructure for RSI includes GPUs and sandboxes, it also has to include the measurement side: analog signal detection with lower noise floors, and deeper instruments such as high-precision microscopes and radio telescopes, built to scale the way compute does. Without that, RSI stays where verification is cheap and shallow.
Two corners: slow and fragile
Two workloads I care about fail on different axes.
AutoML / MLE is slow. Each full candidate evaluation includes training. The verdict is reasonably meaningful when the data protocol is sound, but you get few of them. The AIRA₂ paper, Figure 4, is a useful data point on precision here: their experiments with a fixed hidden split protocol suggest that “the degradation in prior work was due to evaluation noise, not true data overfitting.” This is evidence about the studied setup; the authors explicitly leave open whether true overfitting emerges with more compute.
End-to-end megakernels are fragile. A candidate runs quickly, but correctness is a numerical tolerance against a reference on some set of inputs, random inputs miss edge cases, bf16 makes the tolerance a judgment call, and timing is noisy. Worse, a search or RL loop optimizing against the benchmark can find holes in it. In a February 2025 update, Sakana AI acknowledged that AI CUDA Engineer had found ways to exploit its verification sandbox, and its follow-up report documents benchmark loopholes that inflated the measured speedups. A megakernel adds an end-to-end question on top: does the fused kernel still produce the same model output over a full decode, not just the same per-op tensors?
So one workload is bottlenecked on speed and the other on precision. Infrastructure for RSI has to deliver both.
What the infrastructure should look like
A fidelity ladder with caching at every rung
| Rung | AutoML | Megakernel | Cache key |
|---|---|---|---|
| L0 static | imports, shape inference, data-split fingerprint check | compiles, register/smem budget, sync structure | code hash + environment hash |
| L1 smoke | a few steps on a data shard, loss decreasing | per-op numerics vs. reference on small shapes | + input hash; golden outputs cached |
| L2 low fidelity | inherit parent weights, train \(k\) steps, paired control | boundary graph replay of the edited region; microbenchmarks on representative shapes, interleaved A/B timing | + parent checkpoint / build artifact hash / boundary tensor hash |
| L3 ground truth | top-\(k\) retrained from scratch | full model: logit tolerance, greedy-decode token agreement, end-to-end tokens/s | + hardware / driver version |
| Calibration | Kendall \(\tau\) between lower rungs and L3 on a sample | same | not cached, run periodically |
Candidates are promoted rung by rung, as in Hyperband and its asynchronous variant ASHA. Lower rungs only prune. Anything that becomes a reward signal for a learning proposer comes from L3, or from a lower rung that has been calibrated against L3.
Calibration is measured with Kendall’s \(\tau\), a rank correlation. Take a sample of candidates, score each one on a cheap rung and on L3, and look at every pair: the pair is concordant if both rungs order it the same way. Then \(\tau\) is the number of concordant pairs minus discordant pairs, divided by the number of pairs. At \(\tau = 1\) the cheap rung ranks candidates exactly as ground truth does; near \(0\) its ranking carries no information. Because a search only needs the order of candidates, not their exact scores, \(\tau\) is the right test for whether a cheap rung can stand in for L3.
Content-address everything
Code, data splits, checkpoints, compiled artifacts, golden outputs, activation caches. A cache hit should be a lookup on content, not a judgment call. For checkpoints, the git object model applied to tensors is enough.
Incremental verification rests on a locality assumption
Both workloads want to reuse work across a parent and its child, and both break in the same way.
- In AutoML, inheriting the parent’s weights is only valid if the data split and label processing did not change. If the parent trained on rows that are now in the child’s validation set, the inherited weights are leakage and the fitness is inflated. The inheritance key has to include a fingerprint of (dataset version, split, label transform).
- In a megakernel, “I only changed one region, so the rest is still verified” is not guaranteed. Fusion couples regions through synchronization, shared-memory allocation, and scheduling.
So a cache hit has to depend on whether a change crosses the locality boundary, not only on the code diff. When it crosses, go back to L3.
Boundary graph replay for megakernels
Locality can also be made into a tool. Run the parent megakernel once on a real workload and record the tensors that cross the boundary of each region of its task graph. When an edit touches region \(R\), replay only \(R\) on its recorded inputs and compare against its recorded outputs.
This buys three things. The check is cheap, because it runs one region instead of a full decode. A failure is localized to the region that was edited. And the inputs are real activations from a real workload rather than random tensors, so they exercise value ranges and edge cases that random inputs miss. The cache key is (region code hash, boundary tensor hash), so the recording is shared by every child that edits the same region of the same parent.
Step 3 is the drift problem, and it is the main reason end-to-end megakernel verification is hard. A fused kernel rarely reproduces the reference bit for bit: bf16 rounding, a different reduction order, or a different tiling changes the last bits of each region’s output. In a full decode those differences do not stay local. Each region’s output is the next region’s input, each layer feeds the next, and each generated token feeds the next step, so a tiny per-region error compounds until a logit margin is crossed and the greedy token flips. After that the two decodes are no longer comparable token by token. Replay helps in a specific way: because every region’s inputs are reset to the recorded reference, it measures each region’s own error with the compounding removed. That turns “the outputs drifted somewhere” into “this region added more error than its budget”, which is the question an edit-level verifier can answer. Whether the accumulated drift stays inside the budget is still a property of the whole decode, so the end-to-end check remains, and replay tells you where to look when it fails.
Step 4 is the other limit. Replay checks the dataflow semantics of \(R\) in isolation. It does not see what the edit does to neighboring regions through synchronization, shared-memory aliasing, or scheduling, and an isolated timing says little about end-to-end throughput when regions overlap. So replay belongs on L1/L2, and the full-model check on L3 still has to run, at a frequency set by calibration.
Weight sharing is the AutoML lever worth exploring
In AutoML the dominant cost of a verification is training. Cascades and early stopping save money by discarding candidates sooner; they do not make a surviving candidate cheaper to evaluate. Weight sharing is the one lever that attacks the training cost itself.
Neural architecture search ran into this first. Zoph and Le trained every sampled architecture from scratch, and so did regularized evolution. ENAS made search far cheaper by letting all candidates share the weights of one over-parameterized supernet. A supernet does not fit an RSI loop, because it assumes a fixed, enumerable search space, and an LLM operator edits arbitrary code. What does fit is lineage inheritance, which predates supernets: in large-scale evolution, children start from their parent’s weights; Net2Net grows a network while preserving its function; LEMONADE combines such morphisms with Lamarckian inheritance; and Population Based Training copies weights between population members during training.
Making this work in an open program space needs an inheritance resolver that decides, module by module, what a child may reuse:
The rules behind the diagram:
- Match, then copy with optimizer state. Match parameters by module path, shape, and a hash of the module’s normalized source. Inheriting weights without their optimizer state makes the first steps noisy enough to ruin a short-run comparison.
- New modules start as zero-init residual branches, so the child computes the parent’s function at step zero. Do this as an automatic post-processing step rather than trusting the LLM to follow the convention.
- Crossover by module, not by averaging. Take each module whole from one parent. Two independently trained networks differ by a permutation of hidden units, and Git Re-Basin shows that weight-space merging only works after aligning them.
- A data fingerprint in the inheritance key, so inheritance never smuggles validation data into training.
- An activation cache. Many MLE solutions are a frozen pretrained backbone plus a head, where the expensive part is the backbone’s forward pass. Gradient-boosted trees get nothing from either inheritance or caching, so the share of neural solutions in the search tree bounds the total benefit.
- Compute accounting. Each node carries cumulative FLOPs, inherited plus its own. Old lineages accumulate training and will otherwise crowd out fresh drafts.
The open question is ranking fidelity. Yu et al. showed that weight sharing can scramble the ranking of candidates, and that this matters because the ranking is often used to train the sampler. In an RSI loop the problem is sharper, because fitness is both the selection signal and the reward for the proposer.
Paired comparisons raise information per verification
A child that inherits weights and trains \(k\) more steps will usually beat its parent simply because it trained longer, so a learning proposer is rewarded for near-no-op edits. The cheapest fix I know is a null-edit sibling: for each parent, also run “same code, \(k\) more steps” with the same LR schedule and optimizer state, and score children against that control. One control per parent, shared by all its children.
The kernel equivalent is interleaved A/B timing on the same GPU, which cancels thermal and clock drift.
Reward the best child, then re-verify it
What the reward measures matters as much as how it is measured. TTT-Discover makes the case for discovery problems: “naive RL optimizes average performance, and is indifferent to the state of the art. In discovery, however, success is determined by the maximum.” It replaces the expected reward with an entropic objective,
\[J_\beta(\theta) = \mathbb{E}_{s}\left[\log \mathbb{E}_{a \sim \pi_\theta(\cdot \mid s)}\, e^{\beta(s) R(s,a)}\right],\]which tends to the maximum as \(\beta \to \infty\), and when it picks which state to expand next it scores a state by the maximum reward among its children rather than their mean. An RSI loop has the same shape: the archive only moves when one child beats the frontier, so the proposer should be paid for the best child, not the average one.
A max-seeking reward makes verification noise more dangerous, not less. The maximum of many noisy measurements is biased upward: with enough children, one of them will look like an improvement by chance, and an objective that rewards the maximum will find it. So the two ideas go together. Reward the best child, then re-verify that child under a paired control before its score enters the archive or the gradient. The cost of the second check is paid once per winner, not once per child.
Treat the verifier as an adversarial target
If the verifier’s output is optimized against, the verifier needs version control, regression tests, and red-teaming like any other security boundary. AIRA₂’s hidden evaluation protocol, where agents observe scores computed by the orchestrator in a separate container while search and selection labels remain hidden, is one concrete version of this.
Schedule heterogeneous verification
A verification queue that mixes ten-second smoke tests with ten-hour training runs needs preemption and priorities, or the cheap rungs starve behind the expensive ones and the ladder loses its point. Synchronous batches add a second problem: they wait for their slowest member, so long-tailed verification pushes toward asynchronous updates.
So, is the infra ready?
Not yet. The loop can make verification itself an object of improvement: cheaper proxies, inheritance, replay, early stopping, cascades. But to stay verified, every proxy needs periodic ground truth, and that frequency can go down but not to zero. The real lower bound is the ground-truth rate needed to keep proxy error under a threshold, and that number is measurable: it is the fraction of from-scratch runs needed to keep \(\tau\) above your cutoff.
What this post concludes:
- Verification sets the clock. Once the target involves training, faster generation barely moves the loop. Speed up the side that dominates the iteration, and for RSI that side is verification.
- Verifiability is partly an infrastructure property. Each of verifier’s rule’s five properties has an infrastructure lever, so a task can be made more verifiable without changing the task.
- Reuse work, but only inside locality boundaries. Weight inheritance for AutoML, boundary graph replay for megakernels, and content-addressed caches make most checks cheap. An edit that crosses a boundary, such as a new data split or a shared-memory change, has to go back to ground truth.
- Noise, not compute, sets the depth. Paired controls lower the cost of a reliable verdict, but nothing changes its \(1/\delta^2\) slope.
- Rewards come from calibrated signals. Cheap rungs prune; ground truth, or a rung calibrated against it, pays the proposer. Reward the best child, as discovery requires, but re-verify that winner before it counts. The ground-truth rate can go down but never to zero.
- The verifier is an attack surface. Anything optimized against it needs hidden labels, versioning, and red-teaming.
- Physical RSI needs scalable instruments. Without lower noise floors and deeper, more parallel measurement, RSI stays where verification is cheap and shallow.
In the terms of verifier’s rule: the task RSI needs, improving ML systems, is not very verifiable today, and most of what would make it more verifiable is infrastructure. Beyond code, the same holds for instruments: an RSI loop that reaches into the physical world is only as deep as its measurements are precise. My bet is that the systems that matter for RSI will not be the ones that generate the most candidates. They will be the ones with a fast, precise, cached, and calibrated verification service underneath.
References
- Wei. Asymmetry of verification and verifier’s rule 2025.
- Romera-Paredes et al. Mathematical discoveries from program search with large language models (FunSearch). Nature, 2024.
- Novikov et al. AlphaEvolve: A coding agent for scientific and algorithmic discovery 2025.
- Wang et al. ThetaEvolve: Test-time Learning on Open Problems ICML 2026.
- Yuksekgonul et al. Learning to Discover at Test Time (TTT-Discover). 2026.
- Yang et al. Frontis-MA1: Training an AI4AI Model towards Recursive Self-Improvement in Machine Learning Engineering 2026.
- FrontisAI. OpenRSI: OpenMLE-Gym, OpenMLE-RL, OpenMLE-Evo, and Frontis-MA1
- Zoph, Le. Neural Architecture Search with Reinforcement Learning ICLR 2017.
- Real et al. Large-Scale Evolution of Image Classifiers ICML 2017.
- Real, Aggarwal, Huang, Le. Regularized Evolution for Image Classifier Architecture Search AAAI 2019.
- Pham et al. Efficient Neural Architecture Search via Parameter Sharing (ENAS). ICML 2018.
- Chen, Goodfellow, Shlens. Net2Net: Accelerating Learning via Knowledge Transfer ICLR 2016.
- Elsken, Metzen, Hutter. Efficient Multi-objective Neural Architecture Search via Lamarckian Evolution (LEMONADE). ICLR 2019.
- Jaderberg et al. Population Based Training of Neural Networks 2017.
- Yu, Sciuto, Jaggi, Musat, Salzmann. Evaluating the Search Phase of Neural Architecture Search ICLR 2020.
- Ainsworth, Hayase, Srinivasa. Git Re-Basin: Merging Models modulo Permutation Symmetries ICLR 2023.
- Li et al. Hyperband: A Novel Bandit-Based Approach to Hyperparameter Optimization JMLR 2018.
- Li et al. A System for Massively Parallel Hyperparameter Tuning (ASHA). MLSys 2020.
- Hambardzumyan et al. AIRA₂: Overcoming Bottlenecks in AI Research Agents 2026.
- Lange et al. Towards Robust Agentic CUDA Kernel Benchmarking, Verification, and Optimization 2025.
- Sakana AI. AI CUDA Engineer evaluation update. February 2025.
- Jain. Needle in a haystack ATLAS blog, 2012.
- ATLAS Collaboration. Observation of a new particle in the search for the Standard Model Higgs boson with the ATLAS detector at the LHC Physics Letters B, 2012.
- CMS Collaboration. Observation of a new boson at a mass of 125 GeV with the CMS experiment at the LHC Physics Letters B, 2012.
- Goddi et al. First M87 Event Horizon Telescope Results and the Role of ALMA The Messenger, 2019.
- Zadouri et al. FlashAttention-4: Algorithm and Kernel Pipelining Co-Design for Asymmetric Hardware Scaling 2026.
- NVIDIA. Pushing Latency Boundaries: Optimizing DeepSeek-R1 Performance on NVIDIA B200 GPUs TensorRT-LLM tech blog.
- Cemri et al. SPECS: Faster Test-Time Scaling through Speculative Drafts 2025.
- Gat et al. Discrete Flow Matching NeurIPS 2024.
- Ruan et al. SpecRover: Code Intent Extraction via LLMs ICSE 2025.
- Karpathy. Reproducing GPT-2 (124M) in llm.c in 90 minutes for $20 2024.
- GoedelMachines. Goedel-mHC-1B model card.
- Lambda. Open model, open metrics: How Lambda and the Olmo team trained Olmo Hybrid
- Touvron et al. LLaMA: Open and Efficient Foundation Language Models 2023.
- Hazy Research. Look Ma, No Bubbles! Designing a Low-Latency Megakernel for Llama-1B 2025.