Skip to main content
NVIDIA
AI Factory Operations Agent UI
Container
NVIDIA
AI Factory Operations Agent UI

Web interface, API, CLI, and governed MCP gateway for evidence-backed AI factory operations.

Subscribe to get accessSubscribe to the product below to access this premium content:
NVIDIA Mission Control
NVIDIA Mission ControlNVIDIA Mission Control™ powers every aspect of AI factory operations — from developer workloads to infrastructure to facilities — with the skills of a world-class operations team delivered as software. It powers NVIDIA Blackwell™ data centers for the newest frontiers of AI, bringing instant agility to inference and training workloads and full-stack intelligence that delivers world-class infrastructure resiliency. Mission Control lets every enterprise run AI with hyperscale-grade efficiency so you can accelerate AI experimentation.
Note: You can gain access to hundreds more GPU-optimized artifacts by creating a free NGC account.
Already Subscribed?Log in
Subscribe Now

NVIDIA AI Factory Operations Agent UI

This container provides the web interface, API, command-line interface, and Model Context Protocol (MCP) gateway for the NVIDIA AI Factory Operations Agent blueprint. It gives infrastructure operators and site reliability engineers a single interface for evidence-backed investigation of AI factory cluster health and workload failures.

Capabilities

  • Conversational investigation across Kubernetes, Slurm, Base Command Manager, Prometheus, and configured operational systems.
  • Access to specialized diagnostic and research agents deployed by the blueprint.
  • Evidence presentation, investigation history, and downloadable diagnostic results.
  • Governed MCP tool execution with explicit controls for operations that can modify infrastructure.
  • Administration of supported agent connections and Research Agent document collections.

Intended use

Use this container as the user-facing service in an AI Factory Operations Agent deployment. Typical workflows include investigating unhealthy nodes or workloads, correlating infrastructure telemetry, retrieving operational documentation, and reviewing recommended remediation steps. It is an experimental operations aid and does not replace an organization's monitoring, incident-management, access-control, or change-management systems.

Deployment requirements

Deploy this image through the AI Factory Operations Agent Helm chart. The complete deployment requires a Kubernetes cluster, access to a supported OpenAI-compatible language-model endpoint, and connectivity to the operational services enabled in the chart. Individual integrations can require additional service credentials and Kubernetes RBAC configuration. The UI container itself does not require a GPU. Images are provided for Linux AMD64 and ARM64.

Installation, configuration, architecture, security controls, and support policy are documented in the AI Factory Operations Agent repository.

Publisher

Built and published by NVIDIA Corporation. Licensed under the Apache License 2.0.

Publisher
NVIDIA
Latest Tag0.0.1
UpdatedSeptember 18, 2026 UTC
Compressed Size595.89 MB
Multinode SupportNo
Multi-Arch SupportYes

NVIDIA uses cookies to improve your experience on our web site. We and our third-party partners also use cookies and other tools to collect and record information you provide as well as information about your interactions with our websites for performance improvement, analytics, and to assist in marketing efforts. By clicking "Accept All", you consent to our use of cookies and other tools as described in our Cookie Policy. You can manage your cookie settings by clicking on "Manage Settings." By continuing to use this site or by clicking one of the buttons below, you agree to our Terms of Service (which contains important waivers). Please see our Privacy Policy for more information on our privacy practices.