Skip to main content
NVIDIA
NVOneFormer Commercial
Model
NVIDIA
NVOneFormer Commercial

OneFormer - a unified AI model for multiple image segmentation tasks - trained on commercial data.

This model is backed by NVIDIA's Plus Plus (++) Promise
NVIDIA's ++ Promise covers the quality of the datasets used to train this model.
FieldResponse
Intended Task/Domain:Segmentation (specifically instance, semantic, and panoptic)
Model Type:Transformer (Masked-attention Mask Transformer)
Intended Users:Generative AI creators working with conversational AI models and image content.
Output:Label, Mask and Score for each detected object in the input image.
Describe how the model works:The model takes an RGB image as input. A backbone feature extractor (like a Swin Transformer) creates image features. These features are fed into a transformer decoder that uses masked attention. This attention mechanism extracts localized features by constraining cross-attention within predicted mask regions, allowing it to output a set of segmentation masks and their associated class labels and confidence scores.
Name the adversely impacted groups this has been tested to deliver comparable outcomes regardless of:Not Applicable
Technical Limitations & Mitigation:This model is designed for use with the NVIDIA TAO Toolkit and requires an NVIDIA GPU with sufficient memory (e.g., >12GB). Like many segmentation models, its performance may degrade on object classes that were under-represented in its training data (e.g., unusual objects or scenarios). Mitigation involves fine-tuning the model on a custom dataset that includes more examples of these under-represented classes.
Verified to have met prescribed NVIDIA quality standards:Yes
Performance Metrics:mean Intersection over Union (mIoU)
Potential Known Risks:The model may inaccurately segment objects or fail to detect them entirely, especially if they are small, heavily occluded, or belong to a class not well-represented in the training data. This could lead to incorrect object counting, identification, or area measurement in a downstream application.
Licensing:Use of this model is governed by the NVIDIA Open Model Agreement

NVIDIA uses cookies to improve your experience on our web site. We and our third-party partners also use cookies and other tools to collect and record information you provide as well as information about your interactions with our websites for performance improvement, analytics, and to assist in marketing efforts. By clicking "Accept All", you consent to our use of cookies and other tools as described in our Cookie Policy. You can manage your cookie settings by clicking on "Manage Settings." By continuing to use this site or by clicking one of the buttons below, you agree to our Terms of Service (which contains important waivers). Please see our Privacy Policy for more information on our privacy practices.