AI can inspect more code than a human security team can review manually, but volume alone does not create safer software. A scanner that generates persuasive false positives will be ignored; an autonomous patcher without containment can introduce new risk. The production pattern is a layered system in which agents discover possibilities and deterministic controls decide what may advance.
Google has described an agentic security pipeline that performs lightweight pre-submit scanning, validates findings against code structure, adds deeper post-submit analysis, and proposes fixes for human review. Its open-source Mantis work provides another useful principle: exploratory agents should sit inside a programmatic harness that enforces state, sandboxing, and stage boundaries.
Separate discovery from proof
A discovery agent is useful for semantic issues that signature rules miss: authorization gaps, unsafe assumptions, and multi-step exploit paths. Its output is a hypothesis. A validation stage should then inspect abstract syntax, call graphs, data flow, configuration, and reachability to determine whether an attacker can exercise the path.
Keeping the stages separate reduces confirmation bias. Give the validator independent instructions and evidence. Use existing SAST tools as high-signal inputs rather than asking the model to rediscover every known rule.
Pin the code under review
Long-running analysis becomes unreliable if the target changes mid-pass. Record the commit or immutable snapshot used for discovery, validation, reproduction, and patching. A finding against one version should not silently be tested against another.
This also improves auditability: reviewers can reproduce the exact source, tool versions, evidence, and proposed change that produced a verdict.
Contain dynamic verification
Reproduction may execute untrusted code or agent-generated payloads. Run it in an isolated environment with no production credentials, restricted network access, bounded resources, and disposable state. The agent should never decide to widen its own sandbox.
Define attempt limits and terminal outcomes. Failure to reproduce is not proof of safety; successful reproduction is not permission to test a live system. Any staging or production verification needs an explicit human gate.
Make patching accountable
A fix agent should use the validated exploit path and repository standards, produce a minimal patch, and run deterministic tests. Present the change with the vulnerability evidence, affected behavior, residual risk, and rollback plan. A qualified reviewer remains responsible for acceptance.
Measure precision, adoption, escaped vulnerabilities, review time, regressions, and mean time from detection to verified remediation. High finding counts can indicate noisy automation rather than improved security.
Integrate at two speeds
Pre-submit checks need tight latency and should focus on local changes with strong context. Nightly or scheduled scans can analyze cross-module behavior and larger dependency graphs. Use the same finding identity and lineage so deeper scans enrich rather than duplicate earlier work.
The takeaway
Agentic security works best as a controlled pipeline: discover broadly, validate structurally, reproduce in containment, patch minimally, and require accountable review. The model contributes reasoning; the harness contributes trust.
Sources
- Google Cloud: Using agentic AI to secure infrastructure code
- Google Mantis open-source repository
- Google Mantis: guidance for production agent harnesses
Build it with Cogniquaint experts
Cogniquaint’s AI and security specialists can help your engineering team design a staged review harness, integrate deterministic scanners, isolate reproduction, define approval gates, and measure whether agentic security improves real outcomes.
Work with Cogniquaint
Ready to elevate your operations with AI-powered insights?
Get in touch with us to build your next intelligent solution.




