TermiGen: High-Fidelity Environment and Robust Trajectory Synthesis for Terminal Agents Paper • 2602.07274 • Published Feb 6 • 212
TermiGen: High-Fidelity Environment and Robust Trajectory Synthesis for Terminal Agents Paper • 2602.07274 • Published Feb 6 • 212
VulnLLM-R: Specialized Reasoning LLM with Agent Scaffold for Vulnerability Detection Paper • 2512.07533 • Published Dec 8, 2025 • 4
VulnLLM-R Collection Data and model for VulnLLM-R: Specialized Reasoning LLM with Agent Scaffold for Vulnerability Detection • 9 items • Updated Dec 17, 2025 • 9
SecCodePLT: A Unified Platform for Evaluating the Security of Code GenAI Paper • 2410.11096 • Published Oct 14, 2024 • 13
OWL: Optimized Workforce Learning for General Multi-Agent Assistance in Real-World Task Automation Paper • 2505.23885 • Published May 29, 2025 • 1
AgentVigil: Generic Black-Box Red-teaming for Indirect Prompt Injection against LLM Agents Paper • 2505.05849 • Published May 9, 2025
PromptBench: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts Paper • 2306.04528 • Published Jun 7, 2023 • 3
A Survey on Evaluation of Large Language Models Paper • 2307.03109 • Published Jul 6, 2023 • 43
Improving Generalization of Adversarial Training via Robust Critical Fine-Tuning Paper • 2308.02533 • Published Aug 1, 2023
Large Language Models Understand and Can be Enhanced by Emotional Stimuli Paper • 2307.11760 • Published Jul 14, 2023 • 1
DyVal: Dynamic Evaluation of Large Language Models for Reasoning Tasks Paper • 2309.17167 • Published Sep 29, 2023 • 1