Traditional fuzzers generate random inputs or mutate existing ones. LLM Fuzz CI flips the model: an agent reads your code, understands your test assertions, and generates adversarial inputs designed to violate them. The agent writes inputs, your existing test suite replays them, and only real assertion failures get reported. No static analysis warnings, no CVE spam, no false positives.
The tool runs as a GitHub Action. You mark tests with a decorator, set a spend limit, and the agent tries to break your assertions. Every reported bug is a real failure of your own tests.
How the Agent Generates Test Cases
The agent receives three things: your source code, the test you marked, and the assertion logic inside that test. It does not generate new assertions. It generates inputs that it believes will cause your existing assertions to fail.
The workflow:
- Developer marks a test with
@pytest.mark.llm_fuzz(budget_usd=5, params=["user_input"])or the vitest equivalent - Agent reads the test file and the code under test
- Agent generates adversarial inputs targeting the specified parameters
- CI runs the normal test suite with the generated inputs
- Only failing tests get reported
The agent does not run the tests itself. It writes inputs to a file, and your standard test runner (pytest or vitest) replays them. This means the agent cannot hallucinate a vulnerability. If the test passes, nothing gets reported.
State Management Between CI Runs
The tool does not persist state between runs by default. Each CI execution starts fresh. The agent does not remember previous findings or learn from past failures.
This design choice avoids complexity but means the agent might rediscover the same bug multiple times. The authors suggest running it on code changes or on a schedule, not on every commit.
For teams that want persistence, you could:
- Store generated inputs in the repo (commit them after each run)
- Use a separate artifact store and pass previous inputs as context
- Build a feedback loop that marks inputs as “already tested” in a database
None of this is built in. The tool is stateless by design.
Feedback Mechanism for Real Vulnerabilities
The feedback loop is simple: if the test fails, it’s a real bug. If the test passes, the input was not adversarial enough.
The agent does not classify severity or decide what counts as a vulnerability. It generates inputs, the test runs, and the assertion either holds or breaks. The developer wrote the assertion, so the developer defined what “correct” means.
This eliminates the false positive problem that plagues static analysis tools. A CVE scanner might flag a dependency as vulnerable even if your code never calls the vulnerable function. LLM Fuzz CI only reports inputs that actually break your tests.
The downside: if your tests have weak assertions, the agent will not find bugs. If you test status_code == 200 but do not check the response body, the agent might generate inputs that corrupt data without failing the test.
CI Integration and Sandbox Boundaries
The tool runs as a GitHub Action. The agent phase and the test execution phase are separate steps.
Agent phase:
- Reads code from the repo
- Calls an LLM API (OpenAI or Anthropic)
- Writes generated inputs to a file
Test phase:
- Runs your normal test suite
- Loads the generated inputs
- Reports failures
The agent does not execute arbitrary code. It writes JSON or Python data structures. The test runner loads those inputs and passes them to your functions.
Security boundaries:
- The agent has read access to your repo (it needs to understand the code)
- The agent does not have write access to your codebase (it only writes test inputs)
- The test runner executes in the same environment as your normal CI tests
- If your tests already run untrusted inputs safely, this adds no new risk
If your tests do not handle untrusted inputs safely, this tool will find that problem. That is the point.
Deployment Shape and Cost Control
The tool is a GitHub Action. You add it to your workflow YAML:
- name: LLM Fuzz CI
uses: tiime-software/llm-fuzz-ci@v1
with:
budget_usd: 10
llm_provider: openai
api_key: ${{ secrets.OPENAI_API_KEY }}
The budget_usd parameter caps spending. The agent stops generating inputs when it hits the limit. This prevents runaway costs if the agent gets stuck in a loop or generates thousands of inputs.
Cost scales with:
- Number of marked tests
- Complexity of the code under test
- How many inputs the agent generates before finding a failure
The authors report spending $5-10 per run on typical codebases. For a daily CI run, that is $150-300 per month. For a run on every PR, costs scale with PR volume.
Failure Modes and Observability
Agent generates useless inputs: If the agent does not understand the code, it might generate inputs that do not exercise interesting code paths. The test passes, nothing gets reported, and you wasted API credits.
Mitigation: start with simple tests and check that the agent generates plausible inputs. If it generates random strings for a function that expects JSON, the agent did not understand the schema.
Agent generates too many inputs: If the agent finds a bug on the first input, great. If it generates 1,000 inputs before finding a bug, you pay for 1,000 LLM calls.
Mitigation: set a low budget for initial runs. Increase it if the agent is not finding bugs.
Test assertions are too weak:
If your test only checks response.status == 200, the agent might generate inputs that corrupt data, leak secrets, or trigger race conditions without failing the test.
Mitigation: write stronger assertions. Check response bodies, database state, and side effects.
Agent finds a bug but the report is unclear: The tool reports which input caused the failure, but it does not explain why the input is adversarial or how to fix the bug.
Mitigation: read the generated input and the test failure. The input is usually small enough to understand manually.
Comparison to Traditional Fuzzing
| Dimension | Traditional Fuzzer | LLM Fuzz CI |
|---|---|---|
| Input generation | Random or mutation-based | LLM reads code and generates targeted inputs |
| Coverage | High code coverage, low semantic coverage | Low code coverage, high semantic coverage |
| False positives | Crashes and timeouts (may not be bugs) | Zero (only reports assertion failures) |
| Cost | CPU time (cheap) | LLM API calls (expensive) |
| Setup | Requires instrumentation and corpus | Requires marking tests |
| Best for | Finding memory corruption, crashes | Finding logic bugs, business rule violations |
Traditional fuzzers excel at finding crashes and memory corruption. LLM Fuzz CI excels at finding logic bugs that violate business rules but do not crash.
Example: a traditional fuzzer might find a buffer overflow in a parser. LLM Fuzz CI might find that a discount code can be applied twice, or that a user can access another user’s data by manipulating a query parameter.
When to Use This
Good fit:
- You have tests with strong assertions
- You want to find logic bugs, not just crashes
- You can afford $5-10 per run
- You trust your test suite to define correct behavior
Bad fit:
- Your tests only check status codes or basic output
- You need exhaustive code coverage
- You are fuzzing C code for memory corruption
- You do not have a test suite
The tool is not a replacement for static analysis, dependency scanning, or traditional fuzzing. It is a complement. Use it to find bugs that other tools miss because they do not understand your business logic.
Technical Verdict
LLM Fuzz CI solves the false positive problem by only reporting inputs that fail your own tests. This makes it useful for finding logic bugs that static analysis tools miss. The cost is manageable for daily or per-PR runs on small to medium codebases.
Use it if you have a strong test suite and want to find adversarial inputs that violate your assertions. Avoid it if your tests are weak or if you need exhaustive code coverage. The tool is only as good as your tests.
The stateless design is both a strength and a limitation. It keeps the tool simple but means you might rediscover the same bugs. For teams that want persistence, you will need to build your own state management layer.
The biggest risk is cost runaway if the agent generates thousands of inputs without finding bugs. Start with low budgets and monitor spending.