top of page

Superagent launches Security-One 27B open model detecting 599 of 600 prompt-injection attacks

2 minutes ago
8 min read
Superagent launches Security-One 27B open model detecting 599 of 600 prompt-injection attacks

Superagent has launched Security-One 27B, an open-weight AI model designed specifically for continuous security decisions across AI agents, applications, code and infrastructure.


Unlike a conventional language model, Security-One is not intended primarily to generate answers. It receives a piece of potentially untrusted information and returns calibrated probabilities for predefined decisions, allowing software to determine whether an event should be allowed, blocked, escalated or sent for human review.


The model contains 27 billion parameters, is derived from the Qwen3.8 architecture, supports a native context window of 262,144 tokens and is available as Apache 2.0 open weights as well as through Superagent's hosted API.


Its headline result comes from Superagent's frozen BIPIA prompt-injection evaluation: at an unsafe-probability threshold of 0.70, Security-One detected 599 of 600 attacks while incorrectly flagging 1 of 200 benign inputs.


That corresponds to 99.83% attack detection and a 0.50% benign false-positive rate on that specific evaluation.


··········


SECURITY-ONE 27B AT A GLANCE


........


Specification

Security-One 27B

Developer

Superagent

Parameters

27B

Architecture

Qwen3.8

Base lineage

Qwen3.8-27B / AutoJev-27B

Primary function

Security decision model

Native context

262,144 tokens

Validated classification context

65,536 tokens

Decision options

2–16

Output

Calibrated probability per option

Prompt-injection threshold

0.70 unsafe probability

Open weights

Yes

Weight format

Merged BF16 SafeTensors

License

Apache 2.0

Hosted API

Security-One API

Self-hosting

Supported


........


The output architecture is the defining difference.


Security-One does not need to write a paragraph explaining whether something appears dangerous. It can instead assign probabilities directly to structured alternatives and let application code enforce the corresponding policy.


··········


SECURITY-ONE DETECTED 599 OF 600 BIPIA ATTACKS


Superagent evaluated the model against a frozen binary adaptation of BIPIA, a benchmark designed around indirect prompt-injection scenarios.


The evaluation contained 800 examples: 600 attacks and 200 benign inputs.


At Superagent's release threshold of 0.70 unsafe probability, Security-One produced the following results:


........


BIPIA result

Security-One 27B

Total examples

800

Attack examples

600

Attacks detected

599

Missed attacks

1

Benign examples

200

Benign false positives

1

Overall accuracy

99.75%

Attack detection rate

99.83%

Benign false-positive rate

0.50%


........


The numbers are unusually strong, but their scope needs to remain precise.


They are Superagent's own model evaluation on its frozen BIPIA adaptation, not evidence that Security-One detects 99.83% of every prompt-injection attack encountered in production.


Performance changes with the attack distribution, input format, language, threshold and deployment environment.


Superagent explicitly states that prompt injection remains an open problem and that the model should not be treated as a complete security boundary.


··········


THE MODEL RETURNS PROBABILITIES INSTEAD OF FREE-FORM SECURITY ANALYSIS


Security-One is built around classification rather than conventional generative output.


An application provides a shared state — for example a user prompt, retrieved document, tool result, code change or security alert — together with a question and a set of possible decisions.


The model then assigns probabilities to those options.


For a prompt-injection detector, an application could define only two possibilities: Safe and Unsafe.


For operational security triage, the alternatives could instead be Monitor, Escalate, Quarantine or another set of mutually exclusive actions.


Security-One supports between two and 16 options per decision.


This architecture makes the model easier to insert into deterministic software because the surrounding application, rather than the model itself, controls what happens after classification.


The model estimates risk. The application owns the enforcement action.


··········


A 0.70 THRESHOLD DETERMINES WHEN AN INPUT IS BLOCKED


Probabilistic security models require an operating threshold.


Superagent uses 0.70 unsafe probability as the release threshold for its prompt-injection screening example.


An input at or above 0.70 can be blocked immediately.


Superagent also demonstrates a three-band architecture:


........


Unsafe probability

Example action

≥ 0.70

Block

0.30–0.70

Escalate

< 0.30

Allow


........


The middle region is particularly important.


Instead of forcing Security-One to make the final decision when confidence is ambiguous, an application can send uncertain cases to a stronger model or a human security reviewer.


This creates a model-routing architecture for security: inexpensive first-pass classification handles obvious cases, while more computationally expensive analysis is reserved for difficult ones.


The thresholds are not universal constants. Superagent recommends recalibrating them using each organization's own traffic and failure costs.


A lower blocking threshold increases attack recall but also increases false positives. A higher threshold reduces unnecessary blocks while accepting greater risk that malicious inputs will pass through.


··········


SECURITY-ONE CAN SCREEN INFORMATION BEFORE AN AI AGENT READS IT


Prompt injection becomes particularly difficult when agents consume information from external systems.


An agent can receive instructions indirectly through:


retrieved web pages, emails, documents, database records, tool outputs, code repositories, support tickets, or other untrusted content.


A malicious instruction embedded inside one of these sources can attempt to override the agent's original objective, expose secrets or manipulate subsequent tool use.


Security-One is designed to operate before that content reaches the main agent.


The security model receives the untrusted material as data and estimates whether it contains an injection attempt. Application code can then allow, quarantine or escalate the content before exposing it to an agent with potentially consequential permissions.


This creates a separate security layer between external information and the model responsible for acting on it.


··········


THE SAME MODEL CAN CLASSIFY TOOL CALLS, CODE CHANGES AND INFRASTRUCTURE ALERTS


Superagent is not limiting Security-One to prompt injection.


The intended deployment surface includes agent tool calls, malicious instructions, agent activity, logs, code changes, vulnerability findings, severity classification and infrastructure alerts.


That broader design explains why Security-One outputs configurable probabilities rather than using a fixed safe-versus-unsafe classifier.


A developer can define the decision being made for each workflow.


A tool call might be evaluated according to whether it violates a policy.


A code change could be routed according to security risk.


An infrastructure event could be assigned to monitoring or escalation.


An agent action could be checked before permission to execute it is granted.


The same underlying model therefore becomes a general security decision layer, while the application determines the decision taxonomy and enforcement rules.


··········


SECURITY-ONE USES A QWEN3.8-BASED 27B ARCHITECTURE


Security-One is based on a 27B Qwen3.8 architecture, continued from AutoJev-27B.


Superagent subsequently trained the model using a rank-32 LoRA for one short epoch before merging the adaptation into BF16 weights.


The post-training corpus contained 18,106 examples spanning security, reasoning, knowledge, preference and operational decision data.


Security-specific training included BIPIA contexts, prompt-injection examples and difficult benign inputs.


Superagent reports performing an exact-hash audit against 1,230 protected evaluation rows and finding no exact overlaps with the training corpus.


The objective is different from ordinary instruction tuning.


Security-One is being optimized to produce calibrated choices that can be consumed programmatically rather than long natural-language responses.


··········


THE MODEL ALSO SHOWS WHERE PROMPT-INJECTION DETECTION REMAINS DIFFICULT


BIPIA is not the only evaluation reported by Superagent.


Performance drops substantially on other datasets.


........


Evaluation

Security-One 27B result

BIPIA overall accuracy

99.75%

BIPIA attack detection

99.83%

BIPIA benign false positives

0.50%

Deepset overall accuracy

88.79%

Deepset attack detection

78.33%

Deepset benign false positives

0.00%

NotInject benign accuracy

87.61%


........


The contrast is important.


On the Deepset evaluation, Security-One detected 47 of 60 attacks, equivalent to 78.33%, rather than the 599-of-600 performance observed on BIPIA.


NotInject presents another trade-off. Because it contains benign examples, it measures whether a sensitive detector incorrectly rejects legitimate inputs rather than whether attacks are caught.


Security-One achieved 87.61% benign accuracy there, meaning its aggressive security operating point rejects a meaningful proportion of more difficult benign cases.


The evaluations therefore demonstrate why a single headline detection percentage cannot describe the complete behavior of a security model.


··········


THE 262K NATIVE CONTEXT IS LARGER THAN THE VALIDATED SECURITY WINDOW


Security-One inherits a 262,144-token native context window from its underlying architecture.


Superagent currently validates classification behavior up to 65,536 tokens.


These numbers should not be treated as interchangeable.


The architecture may technically process a substantially longer sequence, but the published security behavior and calibration have only been validated over the smaller classification context.


For production security systems, calibration matters because the probability itself controls downstream decisions.


If model behavior changes at extreme context lengths, the same 0.70 threshold may no longer produce the same relationship between attack detection and false positives.


Superagent therefore rejects oversized requests in its hosted API rather than silently truncating them.


That design avoids a dangerous failure mode in which the malicious portion of a document could disappear during hidden truncation while the application assumes the complete input was inspected.


··········


OPEN WEIGHTS ALLOW SECURITY CLASSIFICATION TO STAY INSIDE PRIVATE INFRASTRUCTURE


Security-One is available as Apache 2.0 open weights, allowing organizations to run the model without sending sensitive security telemetry to Superagent's hosted service.


The company provides a self-hosting configuration based on SGLang.


Its validated production setup uses a single NVIDIA B200 for BF16 inference, although other recent high-memory accelerators may work with different memory and concurrency settings.


Local deployment is particularly relevant for a security model because its inputs can contain highly sensitive material:


source code, internal logs, credentials exposed in alerts, private documents, agent traces and infrastructure information.


Keeping classification inside an organization's own environment can prevent those inputs from being transmitted to an external inference provider.


Self-hosting introduces another requirement, however: calibration must be verified again if the deployment changes the model through quantization, inference-engine differences or other modifications.


··········


SECURITY-ONE IS DESIGNED TO BE CHEAPER THAN SENDING EVERYTHING TO A FRONTIER MODEL


The architecture also addresses the economics of continuous AI security.


A production agent can process enormous numbers of messages, documents, tool outputs and actions.


Sending every event to a large frontier model for detailed security reasoning creates both latency and inference cost.


Security-One instead acts as a first-pass decision model.


Clear benign events can proceed immediately.


Clearly suspicious events can be blocked.


Only uncertain cases need to be escalated to a more expensive model or human reviewer.


Superagent's hosted API charges 0.05 credits per one million input tokens, with output tokens free, reflecting the fact that each question produces only a minimal decision output rather than a long generated response.


The economic advantage ultimately depends on traffic patterns and escalation rates, but the architecture is designed around avoiding expensive generative analysis for every event.


··········


SECURITY DECISIONS REMAIN OUTSIDE THE MODEL


Security-One's most important architectural choice may be what it does not control.


The model returns probabilities, but it does not own the application's security policy.


Code determines the threshold.


Code determines whether an event is blocked.


Code determines whether an uncertain case reaches another model.


Code determines whether a tool call requires human approval.


This separation reduces the risk of treating a probabilistic model as the security control plane itself.


Superagent recommends retaining deterministic controls including least-privilege permissions, sandboxing and human approval around consequential actions.


The model can also be confidently wrong.


A high probability is therefore evidence used by the security system, not a cryptographic guarantee that an event is malicious or benign.


··········


SECURITY-ONE SHOWS A DIFFERENT ROLE FOR SPECIALIZED AI MODELS


Most model releases compete on generation: better reasoning, coding, writing, multimodality or agent performance.


Security-One targets a different layer of the AI stack.


As agents receive broader permissions and interact with increasingly untrusted external information, organizations need systems capable of evaluating what those agents are about to read or do.


A specialized decision model can operate continuously because it does not need to generate a full security analysis for every event.


Security-One's strongest BIPIA result — 599 attacks detected out of 600 — demonstrates the potential of that approach, while its weaker performance on other evaluations demonstrates why deterministic controls and broader testing remain necessary.


The open-weight release also makes the architecture reproducible inside private infrastructure rather than restricting it to a proprietary security API.


Security-One therefore represents a model category likely to become increasingly relevant as autonomous agents expand: AI models whose primary job is not to perform the task, but to decide whether another AI system should be allowed to perform it.


··········


FOLLOW US FOR MORE.


DATA STUDIOS


datastudios.org

bottom of page