vulnerability management

AI vulnerability assessment needs evidence, not just answers

October 1, 2026
Research into automated CVSS assessment shows why cybersecurity AI needs domain context, source evidence and transparent reasoning before teams can trust it.

AI vulnerability assessment uses machine learning or large language models to analyze vulnerability information, estimate severity and help analysts decide what deserves attention. The promise is easy to understand. Security teams face more findings than they can manually investigate, so AI reads the reports, assigns the metrics and explains what matters.

The dangerous part is also easy to understand. A confident answer can be wrong.

In vulnerability management, an incorrect assessment can distort remediation schedules, send teams toward the wrong systems or create false confidence around a serious exposure. Accuracy matters, but accuracy alone is not enough. Analysts need to see the evidence behind the decision.

A new paper, Answer Is Cheap, Show Me the Evidence! Augmenting Automated Vulnerability Assessment with Evidence, turns that principle into a concrete technical study. The researchers built EAVA, a specialized framework for assessing CVSS metrics from software vulnerability reports while producing supporting clues and reasoning.

The title is blunt because the problem is real: generating an answer is easier than generating an answer a security professional can verify.

What the EAVA research tested

The researchers collected 6,446 vulnerability reports from 1,986 software projects. These reports were not clean blocks of text. Nearly one quarter contained screenshots, and 69.2 percent included at least one code snippet. Some reports depended on information about the affected project that was not fully explained inside the report itself.

That resembles real vulnerability work. Evidence arrives as descriptions, crash logs, proof-of-concept code, screenshots, repository context and incomplete technical notes. A model that ignores the inconvenient parts may produce a polished but shallow assessment.

EAVA uses specialized agents to analyze screenshots and code, condense those inputs and add project context before assessing CVSS base metrics. The model is trained to provide evidence and reasoning alongside its result.

According to the paper, EAVA improved severity-weighted F1 by 14.4 percent and Matthews correlation coefficient by 35.2 percent over the strongest baseline. The latter improvement is important because MCC is more informative when metric classes are imbalanced.

The researchers also ran a user study with eight security practitioners. Across the cases where EAVA predicted the correct metric value, participants rated its supporting evidence as useful in 96.8 percent of cases.

These results are promising, but the user study is small. The research focuses on CVSS base metrics, not the complete problem of organizational risk. It also does not establish that every explanation generated by an AI system is faithful to the model's actual internal decision process.

The evidence still supports a clear product principle: security AI should expose what it relied on and make verification easier.

Why general-purpose prompting falls short

One of the paper's more useful findings is that advanced general-purpose models struggled with software vulnerability assessment when used through ordinary zero-shot, chain-of-thought or few-shot prompts. The researchers tested Llama 3.3 70B, DeepSeek V3 and GPT-4.1 across several prompting methods. The general models performed substantially worse than specialized approaches.

The problem was not simply model size. Vulnerability assessment requires domain rules, attention to implicit technical clues and a correct interpretation of the affected software's behavior.

Consider Attack Vector. A short report may not state whether exploitation is local or network-based. The answer can depend on what the affected application does and how users interact with it. Without project context, several values may appear plausible.

This is why “ask the AI” is not a complete security architecture. The model needs grounded inputs, task-specific evaluation and a defined way to express uncertainty.

Evidence should be a product object

Many cybersecurity products treat an AI explanation as a paragraph placed under a score. That is presentation, not evidence.

Evidence should be structured and traceable. For a vulnerability decision, the platform should preserve:

  • the source statement or telemetry supporting each conclusion
  • the time at which the evidence was collected
  • the affected asset, software and version
  • the model or rule that interpreted it
  • the confidence level and unresolved ambiguity
  • the remediation action associated with the conclusion

This structure makes the output useful beyond the analyst screen. It supports review, audit, reporting and future model evaluation. When an assessment changes, the organization can see whether the evidence changed or the interpretation changed.

That distinction matters for explainable AI in cybersecurity. A natural-language explanation can sound reasonable while citing nothing. A grounded explanation lets a person inspect the source.

What this means for Vicarius

Vicarius is building AI into a remediation-first vulnerability management platform. The research offers a standard for how those capabilities should evolve.

The Vicarius AI search assistant lets users query their environment in natural language. The strategic opportunity is to make every important answer inspectable. If a user asks which vulnerabilities require action, the response should point back to the affected assets, risk signals, software evidence and available fixes.

The same principle applies to ScriptAI. Generating a remediation script is only the first step. Trust depends on understanding what the script changes, why those changes address the exposure, what prerequisites it assumes and how the result will be validated.

This does not mean forcing users to read a model's entire reasoning trace. It means providing concise evidence that supports the decision and allowing users to inspect deeper details when needed.

A better trust model for AI-assisted remediation

Security teams should not have to choose between full manual analysis and blind automation. A stronger workflow has five stages:

  1. Collect: Gather vulnerability descriptions, asset data, configuration state, exploit intelligence and remediation options.
  2. Interpret: Use specialized models and deterministic rules to analyze the evidence.
  3. Show: Present the conclusion with source-backed evidence, confidence and known gaps.
  4. Act: Recommend or initiate an action within defined permissions and controls.
  5. Verify: Confirm that the action changed the actual exposure and retain the result.

This workflow treats AI as part of a controlled decision system. It does not ask users to trust eloquence.

The Vicarius reports capability can play an important role in the final stage. Evidence should follow the vulnerability from detection through prioritization, remediation and validation, creating a record that security, IT and compliance teams can use.

From explainable AI to accountable remediation

The EAVA paper focuses on vulnerability assessment, but its larger lesson reaches further. AI becomes operationally valuable when it reduces the work required to reach a trustworthy decision.

That requires domain expertise, contextual data and evidence designed for human review. It also requires humility. A system should not hide uncertainty behind a fluent paragraph.

For vulnerability management, the winning AI experience will not be the one that speaks most confidently. It will be the one that helps teams move from evidence to safe remediation with the least avoidable doubt.

Frequently asked questions

What is evidence-backed AI vulnerability assessment?

It is an automated assessment that links each conclusion to inspectable source material, such as vulnerability text, code, asset telemetry or project context. The evidence allows analysts to validate the result instead of accepting an unsupported score.

Can a general-purpose LLM calculate CVSS accurately?

It can assist, but the research found general-purpose prompting materially weaker than a specialized, context-enriched approach. CVSS assessment requires domain rules and technical evidence that may be implicit or distributed across several sources.

Should AI explanations reveal chain-of-thought reasoning?

No. Operational transparency does not require exposing private model reasoning. Products should provide concise source evidence, decision factors, confidence and validation results.

‍

Sagy Kratu

Sr. Product Marketing Manager

Subscribe for more

Get more infosec news and insights.

Related articles

1000+ members

Turn security converstains into remediation actions