Beyond kernels: Letting agents optimize the whole AI job
Published
AI agents can now write and optimize individual GPU kernels. But there is more to a job’s performance than the speed of its kernels. Data transfers, CPU work, and synchronization can also slow it down.
Recently, we have been trying agents to improve the end-to-end performance of AI jobs. We built /who-ate-my-flops, a small plugin for Claude Code and Codex, to provide profiling tools, workload context, and correctness checks.
Performance improvements made with the tool were merged into six open-source repositories: FunASR, FastVideo, gsplat, Ultralytics, Unsloth, and SGLang, with measured speedups of up to 3.6× on the tested workloads. This blog describes what we learned. We hope the tool can complement kernel-design agents and help extend the use of AI agents to end-to-end optimization of AI workloads.
From kernel optimization to whole-job optimization
A kernel-design agent (KDA, K-Search) typically starts with a reference implementation and test inputs. It then follows an optimization loop: use profiling tools to identify bottlenecks, revise the kernel, verify correctness, and measure performance.
The same loop can also be used to optimize a whole training or inference job. But the optimization target is no longer a single operation. Changes may span the codebase, and each job has its own constraints and correctness requirements.
Optimizing a whole job means reasoning about how its parts interact. A faster kernel may expose a data-loading bottleneck; overlapping communication with computation may require changes to scheduling and memory use. Evaluating an optimization requires checking both its effect on end-to-end runtime and whether it preserves the job’s behavior.
The table below summarizes the axes along which whole-job optimization differs from kernel-only optimization.
| Axis | Kernel optimization | Whole-job optimization |
|---|---|---|
| Optimization goal | Reduce kernel latency. | Reduce total job runtime. |
| Code changes | Usually concentrated in an operator implementation. | May span model code, data loading, dependencies, configuration, and kernels. |
| Optimization constraints | Operator semantics, numerical error, supported shapes, and hardware limits. | Model behavior, memory use, maintainability, and compatibility with other callers. |
| Checking for correctness | Compare outputs with a reference using the task’s tolerances. | Check intermediate values across multiple steps. Each check has its own tolerance. |
More comparisons
| Axis | Kernel optimization | Whole-job optimization |
|---|---|---|
| Inputs | Reference implementation, test inputs, target GPU. | Repository, launch command, data, hardware, and task constraints. |
| Profiling | Nsight Compute: kernel details and hardware counters. | Nsight Systems or torch.profiler: CPU–GPU timelines, transfers, and synchronization. |
| Hardware resources | Often an isolated benchmark on one GPU. | CPU work, host memory, data, and one or more GPUs. |
| Feedback to agent | Often a short, repeatable microbenchmark. | May require initialization, compilation, and multiple steps before a change can be evaluated. |
Where a whole job loses time
The patterns below illustrate some common sources of lost time when looking beyond individual kernels.
/who-ate-my-flops
To address these challenges, we built /who-ate-my-flops.
To get started, give the agent a repository, a launch command, and access to an idle GPU environment. Run /who-ate-my-flops:init to agree on the goal and constraints. Then choose /who-ate-my-flops:diagnose to find performance issues, or /who-ate-my-flops:optimize to iteratively improve end-to-end performance.
The walkthrough below shows this setup for YOLO training with Ultralytics.
Get the plugin and setup instructions ↗
How we designed the harness
We rely on the underlying foundation model to reason about performance and propose optimizations. Our harness provides (1) workload context and profiling evidence to ground that reasoning, (2) correctness checks during optimization, and (3) guidance for the agent to ask users questions and align on goals and constraints.
Two key contexts: computation structure and runtime profiling
Context 1: Computation structure
host → device copies · gradient reduction · synchronization points
computation_graph.jsonexecution_schedule.jsonContext 2: Runtime profiling
torch.profiler / Nsight Systems
Time by operation
Cumulative kernel timeBusy / idle by rank
Communication overlap
Computation structure. The agent reads the code to map out the model’s layers, the modules and operations within them, and how computation, data movement, and synchronization are scheduled. It records this context in JSON files. Together, these describe what we expect the job to do.
Runtime profiling. Profiling (torch.profiler, Nsight Systems) exposes what actually happened at runtime. A raw profile can contain millions of events. We provide scripts for common profiling queries so the agent does not have to write them from scratch. These produce views like those above: time by operation, busy and idle time by rank, and communication overlap. The agent can also write custom queries for each job on its own.
Together, these two contexts help the agent build a mental performance model of the job: where time goes and why. It can compare the measured runtime with the computation structure to identify performance problems that either view alone might miss. We also record source locations (file:line) in the computation structure, so the agent can trace runtime symptoms back to the relevant code.
Correctness checks
Checking whether a change preserves an AI job’s behavior is more complex than checking a single kernel’s outputs. So we built a small tool, parity, to help with these checks. It records intermediate values and compares them across runs. For training, these should include inputs, loss, and gradient norms:
for step, batch in enumerate(loader):
# Record an input checksum.
parity.record(f"batch_step{step}", batch.sum())
loss = model(batch)
# Record the forward-pass loss.
parity.record(f"loss_step{step}", loss, rtol=1e-3)
loss.backward()
# Record the gradient norm.
parity.record(f"gradnorm_step{step}", grad_norm(model), rtol=1e-3)
optimizer.step() # Update the model weights.
The checkpoints are saved to baseline.json and optimized.json. Running parity compare baseline.json optimized.json compares matching tags using the declared tolerances. The report shows which checks failed and by how much, helping the agent investigate the cause:
== 3 tag(s) compared, 2 pass, 1 fail ==
FAIL
gradnorm_step0
expected 3.0 actual 3.012
|diff| 0.012 budget 0.003 4x over budget
The comparison fails. The gradient norm exceeds its tolerance; the input checksum and loss checks pass. Investigate the difference before accepting the speedup.
With parity, users only need to reason about what to check and how much numerical error is acceptable. They can leave the performance work to the agent without having to work through all the ML systems details. In practice, the agent can suggest where to add checks, but we sometimes request additional checks for greater confidence in correctness. We fix random seeds and data order, then run the baseline multiple times to measure the natural variation in the recorded values. This helps the agent distinguish natural variation from differences introduced by an optimization.
Asking users questions proactively
Users may forget to mention some requirements. Those without an ML systems background may not know which details matter. So we let the agent proactively ask questions during init to clarify the user’s goals, constraints, and preferences before optimization begins.
For example, the agent asks, “Are you an ML systems engineer, or an algorithm engineer less familiar with systems?” This helps it tailor its explanations to your background. It also asks what changes are allowed: can it adjust configurations, install faster backends, change Python code, or modify Triton/CUDA kernels? Some users prefer smaller changes that are easier to understand, review, and maintain.
In /who-ate-my-flops, we also provide helpers to turn a job into a repeatable benchmark, reduce file sizes and experiment runtime, and track changes and measurements across experiments. As foundation models improve, agents might need less of this guidance. For now, it helps avoid wasted attempts. The implementation details are in the repository.
Optimization case studies
The agent found several kinds of performance problems, including unnecessary data transfers, repeated initialization, and frequent CPU–GPU synchronization. The two case studies below show how profiling evidence led to concrete changes; the following section summarizes speedups and costs across the evaluated workloads.
FastVideo: unnecessary CPU–GPU transfers in MiniMax H3 inference
We ran MiniMax H3 video generation with FastVideo on four B200s. The agent found that much of the video-decoding time went into moving weights between the CPU and GPU. In the rank-0 interval below, these transfers take 5.87 seconds of the 8.87-second decoding stage.
Interactive trace: GPU stream, VAE decoding stage.
Within the video-decoding stage above, gold shows host-to-GPU transfers, purple shows GPU-to-host transfers, and green shows GPU kernels. The decoder loaded 10.4 GB of weights onto the GPU, then copied them back to the CPU to free GPU memory. That return transfer was unnecessary: inference did not change the weights, so the original CPU copy could be reused.
The agent made two changes for both the video and audio VAEs: it kept the CPU copies and released the GPU copies after decoding, avoiding the return transfers; it also pinned the CPU copies to speed up the remaining host-to-GPU transfers. Together, these changes reduced median end-to-end generation time from 15.20 s to 6.34 s on four B200s—a 2.40× speedup. The changes were merged in PR #1867.
FunASR: repeated attention setup in speech recognition fine-tuning
We fine-tuned Fun-ASR-Nano for speech recognition on one B200. Most steps took about 98 ms, but 59 of the 382 steps took about a second each. The baseline trace below compares a normal step with a slow one.
Interactive trace: forward thread on CPU, backward thread on CPU and GPU stream.
In the slow step above, the CPU attention calls span 444 ms in the forward pass and 554 ms in the backward pass, while the GPU stream below shows only brief bursts of work. The agent’s investigation traced this CPU-side delay to cuDNN building attention execution plans for new batch shapes.
The agent selected other SDPA backends to avoid repeated plan setup, removing the roughly 500 ms CPU-side delays seen above. It also found many small kernels and compiled the decoder to reduce kernel launches. In the benchmark reported in PR #3705, these changes reduced mean step time from 243 ms to 68 ms, excluding the first step—a 3.6× speedup. The changes were merged as opt-in settings.
Not every finding required this much investigation. In our own workloads, the agent sometimes found a missing fast kernel library or a performance-related flag that had not been enabled. These requirements may be missing from the installation instructions, so a coding agent can overlook them when setting up the environment. /who-ate-my-flops can catch and fix these setup issues too.
Experiments: speedups and costs
We explored training and inference workloads from repositories spanning several application domains. Here, we highlight six examples whose optimizations were merged upstream. For each workload, we ran init to clarify the goal and constraints, followed by optimize. We used Fable 5.1, and all runs used B200 GPUs within a single node. The table below reports the measured improvements for each workload. Expand a repository row to see the problems found and how they were addressed.
| Repository | Task | Measured speedup | PR |
|---|---|---|---|
| FunASR 20k ★ | Fun-ASR-Nano speech recognition fine-tuning · 1 B200 | 3.6× mean step | #3705 |
| FastVideo 4k ★ | FastH3 Preview video generation · 4 B200s | 2.40× per generation | #1867 |
| gsplat 5k ★ | 3D Gaussian splatting training · 8 B200s | 1.42× steady step | #1062 |
| ultralytics 61k ★ | YOLO26x object detection training on COCO · 1 B200 | 1.25× step; 1.22× epoch | #26180 |
| unsloth 76k ★ | Qwen3.5-9B LoRA fine-tuning · 1 B200 | 1.09× step | #11238 |
| SGLang 36k ★ | Nemotron-3-Super LLM serving · 8 B200s | 1.075× output token throughput | #41223 |
What a run costs
The charts below show the time spent on init and optimize, along with total token usage for these six workloads.
| Repository | Init | Optimization | Tokens |
|---|---|---|---|
| FunASR | 14m08s | 1h53m | 98M |
| FastVideo | 9m36s | 2h39m | 214M |
| gsplat | 6m20s | 2h50m | 173M |
| ultralytics | 6m11s | 1h13m | 115M |
| unsloth | 8m30s | 2h13m | 139M |
| SGLang | 3m31s | 2h12m | 182.8M |
Most recorded tokens were cache reads. The costs shown exclude human review and later PR revisions.
Where human judgment still matters in end-to-end optimization
Correctness. The agent can add correctness checks to the codebase and run them on its own. Users can help ensure the checks cover what matters for the workload and add stricter ones based on their goals and tolerance for risk.
Tradeoffs. A change that makes one workload faster may not suit other uses of the codebase. It may use more memory, break another use case, or make the code harder to maintain. Human review is needed to understand these tradeoffs and intervene when necessary.
From experiments to PRs. The agent might bundle several unrelated optimizations into one branch. Together, they may give the largest speedup, but they also make the PR harder to review. Users still need to decide which changes belong together in a PR, which should stay as local changes, and whether an optimization should be opt-in rather than enabled by default.
Related work
Agents are now widely believed to design standalone kernels very well. And there is increasing attention on end-to-end performance. Almost all of that attention points at one workload: LLM inference serving. InferenceBench gives an agent a model, one bare H100, and two hours to produce a fast OpenAI-compatible server, using whatever stack it chooses. ISO-Bench points the agent at a bottleneck in vLLM or SGLang that a maintainer has already fixed, and scores its patch against theirs. Both are confined to LLM inference, and both measure one ability: optimizing an inference engine.
ASAP optimizes LLM training instead: it searches for sharding configurations for distributed LLM training on TPU clusters. However, it does not modify the codebase, and it does not verify correctness, on the assumption that changing the parallelization does not change output. Its three experiments are scored by whether the proposal matches what a human expert chose, rather than by a measured A/B.
Many open-source projects have built their own skills for internal performance work.
BBuf's collection and the
skills in SGLang's
own repository cover torch.profiler triage, cross-framework benchmarking and more,
with an autonomous loop that ties them together. All of it is for an inference engine, and most
of it is specific to SGLang. NVIDIA ships a similar family inside
TensorRT-LLM: twenty
skills and ten agents covering kernel writing, bottleneck classification, and an
optimization loop with an explicit revert_file step. It makes a great deal of
internal performance practice portable, and it is scoped to TensorRT-LLM. AMD's Hyperloom agent optimizes model inference on AMD Instinct.
Our goal is to make performance agents easy to use across users’ own training and inference jobs.
Try it on your own workload
The plugin is available for Claude Code and Codex. Parity can also be used separately to record and compare values across runs. We’d like to hear where it helps, where it fails, and what still needs human intervention. Give it a try!
References
- Z.ai. GLM built its inference infrastructure.
- OpenAI. How GPT-5.6 fuses frontier intelligence with frontier efficiency.
- NVIDIA. Kernel Design Agents (KDA).
- Shiyi Cao et al. K-Search: LLM Kernel Generation via Co-Evolving Intrinsic World Model.
- Jehyeok Yeon et al. InferenceBench.
- Ayush Nangia et al. ISO-Bench.
- Yuran Ding et al. ASAP.
- BBuf. AI-Infra-Auto-Driven-SKILLS.
- SGLang. Performance skills.
- NVIDIA TensorRT-LLM. Skills and agents.
- AMD. Hyperloom.