NVIDIA
Grounding DINO
Model
NVIDIA
Grounding DINO

Open vocabulary multi-modal object detection model trained on commercial data.

This model is backed by NVIDIA's Plus Plus (++) Promise
to learn more about the quality of the datasets used to train this model.
FieldResponse
Intended Application(s) & Domain(s):Detecting Objects
Model Type:Object Detection
Intended Users:The model is intended for developers that build multimodal object detection application.
Output:Bounding Box, Confidence Scores, and Noun Phrases
Describe how the model works:Extracts features from provided images using provided list of nouns.
Technical Limitations:The model may have difficulties in non-Flickr style data like medical, satellite, and industrial data.
Verified to have met prescribed NVIDIA standards:Yes
Performance Metrics:Mean Average Precision (mAP)
Licensing:NVIDIA Open Model License