Solutions Introducing Shieldstral. August 4, 2026 By Mistral Back to Blog 5 min read Share this post Copy url to clipboard Copied Thinking Summary Shieldstral introduces a 3B open-weights multimodal safety classifier that outperforms models up to 7x its size by framing content moderation as a policy-adaptive question-answering task. Unlike traditional guardrail models, it accepts plain-language policies at inference time, unifying text and image safety evaluation without retraining. Released under Apache 2.0, it delivers calibrated safety scores across diverse benchmarks while running efficiently on a single 16GB NVIDIA GPU. A 3B open-weights, policy-adaptive multimodal safety classifier that matches models up to 7x its size on text safety and sets a new state of the art on multimodal moderation. “Does this content promote violence against a protected group? Is this image safe to show to a minor? Did the assistant refuse the request?” Every product that ships a model needs to answer questions like these — but the right answer depends on the product, the audience, and the moment. …
From the source
Introducing Shieldstral.
mistral.ai
Covered by 3 outlets