Accept 7.4/10 cs.CL trendtoknow-paper-summaries codex-pro/gpt-5.6-luna

KaliBench: A Fine-Grained Benchmark for Cybersecurity Tool Use on Kali Linux with Runtime-Free Verifiable Rewards

Pengfei Li, Naufal Suryanto, Sicheng Zhang, Muzammal Naseer · October 01, 2026 · cs.CL, cs.AI, cs.CR

📌 Highlights

KaliBench isolates a difficult but operationally important capability: translating analyst intent into exact, executable cybersecurity CLI commands. Its main result is that unrestricted command generation remains far from solved, while documentation-aware evaluation and runtime-free post-training provide strong gains.

  • 8,504 verified query-command pairs cover 1,642 tools, 23 capability dimensions, and 5 security phases.
  • No evaluated open-weight model exceeds 42% unrestricted Exact Correct accuracy.
  • Restricted mode raises average Tool Accuracy from 72.0% to 95.2%; Hinted mode raises average Exact Correct from 22.3% to 73.1%.
  • RedSage-K SFT+GRPO reaches a 79.2% average Total Score with an 8B backbone, compared with 80.2% for DeepSeek-V3.2 at 685/37B parameters.
  • The benchmark remains limited to single-turn command generation and documentation-grounded data.

🎯 Introduction

Existing cybersecurity evaluations primarily measure knowledge through multiple-choice, structured-answer, or classification tasks, while newer agentic benchmarks measure end-to-end performance in environments such as CTF challenges. These settings are useful but do not isolate whether a model can select the correct real-world tool and construct a syntactically valid invocation. Kali Linux is an appropriate testbed because practitioners must operate heterogeneous command-line toolchains, and small errors in flag spelling, alias use, argument ordering, or tool interpretation can invalidate execution.

The paper's objective is to evaluate schema-free natural-language-to-CLI translation directly. Unlike conventional function-calling benchmarks, KaliBench does not assume that every tool and parameter can be placed into an explicit JSON schema. It instead measures whether a model can infer the executable tool and arguments from the request, then decomposes failures into tool selection, optional arguments, positional arguments, and exact command correctness.

🔬 Methodology

The dataset begins with 2,372 Kali tools described in structured official Markdown. For each tool, Qwen3-Max generates 10–13 natural-language query-command pairs with structured annotations. The authors report that more than 50% of initially generated commands contained hallucinated or misspelled options, motivating LLM verification, sandboxed execution, human review, and semantic deduplication. The final dataset contains 8,504 verified pairs; the evaluation split contains 5,000 examples covering all 1,642 tools, and the training split contains 3,504 examples covering 962 tools.

Evaluation uses three information conditions. Unrestricted mode supplies only the query. Restricted mode adds a randomly sampled list of 20 candidate tools containing the correct tool. Hinted mode adds the official documentation of those candidates. Outputs are parsed with Python shlex; optional arguments are alias-aware and order-flexible, while positional arguments must appear in the documented order. The metric set includes Tool Accuracy, Optional-Argument F1, Positional-Argument F1, Total Score, and Exact Correct.

For training, the 3,504 training pairs are expanded across the three prompt modes to create 10,512 samples. LoRA-based SFT, GRPO-based RLVR, and SFT followed by GRPO are trained from the RedSage-Ins 8B backbone. RLVR derives rewards from component-level predictions, exact command matches, and formatting, allowing reward computation without executing model outputs.

Command decomposition: tool

Formula:
\[\texttt{Tool\_name: nmap}\]

Meaning: Identifies the selected Kali executable and isolates tool-selection accuracy from argument construction.

Command decomposition: optional arguments

Formula:
\[\texttt{Optional\_args: \{"--open": null, "-sT": null, "-p": "21-25,80"\}}\]

Meaning: Represents alias-aware flags and their values, enabling component-level matching of optional arguments.

Command decomposition: positional arguments

Formula:
\[\texttt{Positional\_args: ["192.168.1.100"]}\]

Meaning: Preserves ordered non-flag arguments, whose positions are scored separately from optional flags.

Schema-style invocation example

Formula:
\[\texttt{nmap(scan\_type="TCP connect", ports="21-25,80", open\_only=true, target="192.168.1.100")}\]

Meaning: Shows how the same natural-language request can be represented as a structured tool invocation while KaliBench evaluates the executable CLI form.

📊 Experiments

The central benchmark is KaliBench's 5,000-example evaluation set, drawn from the 8,504-pair verified corpus. The authors evaluate 14 general-purpose and 7 cybersecurity-oriented open-weight model configurations, including compact, mid-size, and large-scale models, under Unrestricted, Restricted, and Hinted modes. Open-source inference uses vLLM on NVIDIA A100 hardware; selected large models are evaluated through external APIs. The paper reports Exact Correct, Tool Accuracy, Optional-Argument F1, Positional-Argument F1, and Total Score, where higher values indicate better performance.

The mode comparison shows that tool selection and argument construction are distinct bottlenecks. Restricting the candidate tool list raises average Tool Accuracy from 72.0% to 95.2% but yields only modest gains in argument metrics and Exact Correct. Supplying documentation produces the largest improvement: average Optional-Argument F1 rises from 45.1% to 87.8%, Positional-Argument F1 from 61.1% to 90.1%, Total Score from 59.4% to 90.8%, and Exact Correct from 22.3% to 73.1%.

Model scale alone is insufficient. The strongest open-weight unrestricted Exact Correct result is 41.3% for GLM-5.2, while the compact RedSage-K SFT+GRPO model reaches 32.2% unrestricted Exact Correct and 79.2% average Total Score. The three RedSage-K variants obtain average Total Scores of 76.9%, 77.4%, and 79.2%; SFT+GRPO ranks third among evaluated open-weight models and is reported as only 1.0 percentage point below DeepSeek-V3.2's 80.2%. In the proprietary comparison, GPT-5.6-Sol reaches 61.68% full-set unrestricted Exact Correct, Codex reaches 51.68%, and Claude Opus 5 reaches 44.02%, with response coverage affecting the full-set scores.

Robustness tests on eight model configurations show that paraphrases have little effect, informal queries produce small declines, and messy queries cause the largest drops: -3.20 percentage points in Unrestricted, -1.40 in Restricted, and -1.07 in Hinted. Across capability dimensions, gpu and crypto-stego are relatively strong, while 802-11 and wireless show lower minima; the authors also identify alias misuse, option hallucination, hallucinated long forms, and singular-plural option errors as recurring failure patterns.

tab:combined_perf

Caption: Open-weight model query-to-command performance across three modes. Columns report Unrestricted (U), Restricted (R), and Hinted (H). $^ $ indicates thinking-enabled models, and $^ $ denotes cybersecurity-tuned models. Bold/ underline mark the best and second-best within each size group. Models are sorted by Averaged Total Score. MoE parameters use Total/Active (e.g., 685/37B).

Content:
ModelSize (B)EC UEC REC HTA UTA RTA HOpt-F1 UOpt-F1 ROpt-F1 HPos-F1 UPos-F1 RPos-F1 HTS UTS RTS HAvg. TS
Metric Average–22.328.373.172.095.294.645.148.787.861.168.290.159.470.790.873.6
RedSage-K (GRPO)828.133.066.776.593.393.151.753.784.873.977.088.467.474.688.876.9
RedSage-K (SFT)830.634.477.178.195.997.152.855.290.566.968.492.065.973.293.277.4
RedSage-K (SFT+GRPO)832.237.469.477.995.092.556.157.986.576.980.989.170.377.989.379.2
DeepSeek-V3.2†685/3733.040.284.575.898.397.157.062.093.568.875.494.167.278.694.980.2
GLM-5.2†75341.352.074.886.394.294.975.371.988.451.378.890.671.081.791.381.3

Why It Matters: The table separates tool-selection gains from argument-construction gains. Restricted mode raises average Tool Accuracy from 72.0% to 95.2%, while Hinted mode raises average Exact Correct from 22.3% to 73.1% and average Total Score from 59.4% to 90.8%.

Table 2

Caption: Proprietary models and scaffolded inference in Unrestricted mode.

Content:
ModelAnswered queriesEC (%) answeredEC (%) all 5K
GPT-5.6-Sol4,98861.8361.68
Claude Opus 53,67559.8944.02
Codex (GPT-5.5)4,63355.7751.68

Why It Matters: The table shows that proprietary systems outperform the strongest evaluated open-weight model on full-set unrestricted exact correctness, but coverage and provider safety restrictions materially affect results.

Table 3

Caption: Average change in Total Score, in percentage points, relative to original queries across eight model configurations.

Content:
Query styleUres.Res.Hinted
Paraphrase+0.61+0.22-0.09
Informal-0.55-0.48-0.28
Messy-3.20-1.40-1.07

Why It Matters: The robustness results quantify the cost of noisy language: messy queries cause the largest degradation, while tool documentation reduces but does not remove the decline.

fig:teaser

fig:teaser

Caption: Why KaliBench Matters? KaliBench enables fine-grained, verifiable evaluation and runtime-free reward training for translating analyst intent into executable Kali commands.

Why It Matters: The teaser summarizes the benchmark's central value proposition: decomposing command correctness and using the resulting deterministic signals for model improvement.

fig:kalibench

fig:kalibench

Caption: KaliBench end-to-end pipeline for NL-to-CLI evaluation on Kali Linux. The benchmark is constructed via LLM-based data generation, multi-round verification-regeneration, and representative-sample selection. Evaluation applies alias-aware parsing and fine-grained scoring over tools, arguments, and dimensional categories.

Why It Matters: This is the most useful implementation overview, showing how documentation, verification, deduplication, evaluation, and training fit together.

fig:text visualization

fig:text visualization

Caption: KaliBench Structure and Representative Sample. KaliBench is a structured benchmark of Kali Linux tools organized along the security attack lifecycle. It comprises 5 security phases and 23 functional dimensions. The dataset sample includes natural-language queries and decomposed ground-truth commands. See Appendix text for a detailed view of examples of all 23 dimensions.

Why It Matters: The representative sudo example makes the decomposition target concrete: tool identity, optional aliases and values, ordered positional arguments, and equivalent valid CLI forms.

fig:dim

fig:dim

Caption: Performance across security dimensions and modes. (a) Performance over 23 dimensions and 5 phases per mode. Box plots show distributions (Q1--Q3), with whiskers for min/max. (b,c) Total Score and Exact Correct (\%) across modes; dotted lines indicate inter-mode gaps.

Why It Matters: The figure captures the paper's main empirical pattern: documentation progressively improves performance and reduces variance, while dimension-level difficulty remains.

🔮 Conclusion

The authors conclude that exact cybersecurity CLI generation is substantially harder than general cybersecurity question answering. Restricting the tool search space largely resolves tool-selection errors, but argument-level correctness—particularly optional flags and flag-value bindings—remains the dominant bottleneck. Documentation provides the largest gains, indicating that missing tool knowledge and interface familiarity are major causes of failure.

KaliBench therefore functions as both an evaluation benchmark and a training substrate. Its deterministic decomposition supports fine-grained diagnosis and runtime-free RLVR, while SFT and RLVR together allow an 8B model to approach much larger systems under less-guided conditions.

🛠️ Future Research Improvements

The paper explicitly identifies multi-step workflows, retrieval-augmented tool use, and environment-aware evaluation as future directions. These extensions should preserve KaliBench's component-level diagnostics while adding state tracking, observation interpretation, command feedback, recovery, and authorization checks.

A useful next benchmark version would broaden evaluation beyond documentation-generated queries with naturally occurring analyst requests, controlled ambiguity, and organization-specific command conventions. It should also test whether exact-match improvements survive paraphrase, messy input, unseen tools, and live-but-sandboxed environmental state.

🏭 Potential Industry Use Scenarios

KaliBench can support development of local cybersecurity copilots that translate analyst intent into reviewable command proposals while exposing tool-selection and argument-level confidence. The benchmark's documentation-hinted setting is especially relevant to retrieval-augmented assistants that consult approved tool manuals before producing commands.

Security engineering teams can also use the dataset for regression testing, model selection, and post-training of private models where operational data cannot be sent to external APIs. In production, generated commands should remain proposals subject to authorization, policy validation, target scoping, and sandbox or dry-run checks because benchmark correctness does not establish operational safety.

💬 Critical Analysis

The strongest aspect of KaliBench is attribution: its decomposition distinguishes failure to identify a tool from failure to construct its arguments, a distinction obscured by end-to-end agent benchmarks. Alias-aware parsing, sandbox verification, human review, and runtime-free rewards make the benchmark unusually actionable for engineering teams.

The main validity risk is distributional. The corpus is generated from official documentation, and the paper itself reports that more than 50% of initial generated commands contained hallucinated or misspelled options before filtering. The final verification pipeline improves reliability, but the resulting examples may still underrepresent the ambiguity, shorthand, incomplete context, and organizational conventions present in real analyst requests.

The results also show that exact correctness and component scores can tell different stories. Documentation sharply improves average Total Score, yet exact-command gaps remain wide across models. A builder should therefore treat Total Score as a diagnostic training signal rather than a substitute for end-to-end validation, and should add safety and environment-aware tests before deploying command-generating systems.

Original Abstract

LLMs are increasingly applied to cybersecurity workflows, where they are expected to translate analysts' intent into tool invocations. However, existing evaluations focus on knowledge-based assessments or end-to-end agentic tasks, and do not directly measure LLMs' ability to generate executable commands for real-world cybersecurity tools. This gap is critical because cybersecurity operations rely on strict command-line interfaces (CLIs), where minor syntax errors, incorrect flag--value bindings, or argument misordering can invalidate execution. We introduce KaliBench, a fine-grained benchmark and dataset for natural-language--to--CLI translation on Kali Linux, comprising 8,504 query--command pairs spanning 1,642 tools across 23 capability dimensions and 5 security phases. KaliBench is constructed via a manuscript-grounded pipeline with deterministic canonicalization and alias-aware evaluation, enabling precise and reproducible assessment of tool selection and argument construction. To ensure both semantic correctness and practical executability, we develop a multi-stage verification pipeline that combines LLM-based validation, sandboxed terminal execution, and human-in-the-loop refinement. Building on these fine-grained, deterministic signals, KaliBench further enables runtime-free verifiable rewards for training. Across three evaluation modes and 24 configurations of general-purpose and security-focused open-weight models, no open-weight model exceeds 42% exact-command accuracy in the unrestricted setting, highlighting the difficulty of accurate CLI-based cybersecurity tool use without explicit tool hints. We further show that supervised fine-tuning and reinforcement learning with verifiable rewards derived from KaliBench significantly improve an 8B model and achieve performance comparable to a 685B MoE model.