ai security

Grounded does not mean trusted: securing AI-assisted remediation

October 5, 2026
CodePoisonRAG achieved attack success rates of 80% to 93% with a 0.7% poisoning ratio. Learn what this means for secure AI-assisted remediation.

Retrieval-augmented generation is often presented as an answer to unreliable AI. A RAG system retrieves relevant documents, code or patches and places them inside the prompt. The output is grounded in an external source.

But what happens when the source is malicious?

A new research preprint called CodePoisonRAG examines that exact problem in retrieval-augmented code generation. The researchers created task-specific poisoned artifacts designed to look relevant, retain vulnerable behavior and falsely describe themselves as safe.

They added 85 poisoned artifacts to a corpus of 12,053 entries, an aggregate poisoning ratio of 0.7%. Every artifact appeared among the top three results for its target query. Across three models, attack success ranged from 80% to 93%.

For AI-assisted remediation, the lesson is uncomfortable but useful: grounding can improve relevance without improving trust.

What is knowledge poisoning in retrieval-augmented generation?

Knowledge poisoning is the manipulation of content that an AI system may retrieve and use as context. The attacker does not need to change the foundation model. Instead, the attacker targets the model’s information supply chain.

CodePoisonRAG combines two techniques. First, vulnerability injection adds an attacker-selected source-to-sink weakness while keeping the artifact aligned with the expected programming task. Second, semantic mislabeling adds false safety statements without removing the vulnerable behavior.

The poisoned artifact can look relevant, rank highly and tell the model that appropriate security controls are already present.

The paper’s black-box threat model is also significant. The attacker is assumed to have no access to the victim’s deployed knowledge base, retriever, reranker, generator, prompt or defense. The attack inserts no more than one artifact for each anticipated programming task.

What did the researchers measure?

The team built 85 poisoned artifacts covering ten CWE classes across Java and C. All 85 reached the top three for their target queries. Across Qwen 3.5 9B, Code Llama 13B and DeepSeek-Coder V2 16B, reported attack success ranged from 0.80 to 0.93.

The researchers also tested CodeGuarder, a defense that adds vulnerability-specific security knowledge to the generation context. The attack still achieved success rates between 0.40 and 0.71.

Those results support two evidence-based observations. A very small share of strategically placed content influenced retrieval for every targeted query in the experiment. Additional security guidance reduced the attack, but it did not reliably neutralize poisoned context.

What the study does not establish

CodePoisonRAG was submitted to arXiv on September 2, 2026 and remains under review. It is a controlled research experiment, not a documented compromise of a production remediation platform.

The evaluation covered three models, two languages and ten weakness classes. Some output assessment used an LLM judge guided by CWE-specific rules, and the dataset had not yet been independently reproduced.

These constraints should temper the conclusion. They should not erase it. The paper demonstrates a concrete failure mode under controlled conditions and quantifies how little poisoned content was required.

Why this matters for AI-generated remediation

An AI remediation system can generate a script, configuration change or command that runs across real assets with elevated permissions.

That changes the trust model. A retrieved artifact should not become executable merely because it is relevant, popular or described as secure. Relevance answers, “Does this content match the task?” Trust answers, “Who created it, has it changed, and what evidence shows the action is safe?”

AI-assisted remediation systems should treat every external artifact as untrusted input until controls establish otherwise. Relevant controls include:

  • Provenance that identifies the author, source and review history.
  • Cryptographic integrity checks for approved artifacts.
  • Versioning that makes changes visible and reversible.
  • Static analysis for dangerous commands, insecure patterns and unexpected network access.
  • Sandboxed tests using representative operating systems and application versions.
  • Least-privilege execution with tightly defined asset scope.
  • Human approval for high-impact or novel actions.
  • Post-action validation that confirms the condition changed.
  • Quarantine and revocation when a source or artifact becomes suspect.

These are our defensive implications, not controls evaluated by CodePoisonRAG. The study tested poisoning and one contextual defense. It did not compare a complete remediation governance architecture.

Why semantic mislabeling deserves special attention

Semantic mislabeling pairs vulnerable behavior with text asserting that validation or sanitization is already present.

Those assurances enter the model’s reasoning context and may also reduce a hurried human reviewer’s scrutiny.

A secure review process should evaluate behavior independently from description. Comments, labels, documentation and community votes are useful metadata. None should substitute for analysis of what the script actually does.

What can you the security admin do about it?

Vicarius ScriptAI generates remediation scripts, while vScript supports built-in, custom and community scripts. That makes knowledge provenance and action validation strategically central to the product category.

The Vicarius strongest capability within ScriptAI is not simply that AI can write remediation faster. It is that AI-assisted remediation should connect every action to its source, review status, tested scope, permissions and validation outcome. Speed matters only after trust boundaries are explicit.

Community knowledge creates a software supply chain. Organizations should understand how content enters that chain, how it is tested and what happens if trust is withdrawn.

How should teams evaluate an AI remediation platform?

Buyers should ask direct questions:

1. Which sources can the AI retrieve from?

2. Can administrators restrict retrieval to approved repositories?

3. Is every artifact tied to an author and version?

4. What tests run before deployment?

5. Which permissions does the script receive?

6. Can administrators revoke an artifact everywhere?

7. How does the system confirm that remediation worked?

Clear answers reveal whether AI is attached to an execution engine or embedded in a governed remediation process.

Frequently asked questions

Does RAG make AI-generated code safe?

No. RAG can ground output in retrieved information, but the retrieved information may be vulnerable, outdated or malicious. Source trust and output validation remain necessary.

How much poisoned content did CodePoisonRAG use?

The researchers added 85 poisoned artifacts to a corpus of 12,053 entries, an aggregate poisoning ratio of 0.7%. Each targeted task received at most one task-matched poisoned artifact under the study’s threat model.

Did security guidance stop the attack?

Not completely. Against CodeGuarder, the reported attack success rate remained between 40% and 71%, compared with 80% to 93% across the undefended model evaluations.

My personal takeaway from this research

AI-assisted remediation needs more than accurate retrieval. It needs evidence about the evidence.

The source, integrity, review history and tested behavior of a script should travel with it from generation to deployment. The action should run with limited authority, followed by validation.

Grounding answers the model’s need for context. Governance answers the organization’s need for trust. Secure remediation requires both.

‍

Sagy Kratu

Sr. Product Marketing Manager

Subscribe for more

Get more infosec news and insights.

Related articles

1000+ members

Turn security converstains into remediation actions