NVIDIA NIM for GPU accelerated Qwen2.5-Coder-32B-Instruct inference through OpenAI compatible APIs
Qwen2.5-Coder-32B-Instruct Overview
Description:
This container houses the Qwen2.5-Coder, which is the latest series of Code-Specific Qwen large language models (formerly known as CodeQwen). As of now, Qwen2.5-Coder has covered six mainstream model sizes, 0.5, 1.5, 3, 7, 14, and 32 billion parameters, to meet the needs of different developers. Qwen2.5-Coder brings the following improvements upon CodeQwen1.5:
- Significant improvements in code generation, code reasoning, and code fixing. Increased training tokens to 5.5 trillion including source code, text-code grounding, Synthetic data, etc. Qwen2.5-Coder-32B has become the current state-of-the-art open-source codeLLM, with its coding abilities matching those of GPT-4o.
- A more comprehensive foundation for real-world applications such as Code Agents. Not only enhancing coding capabilities but also maintaining its strengths in mathematics and general competencies.
- Long-context supports up to 32K tokens.
The container components are ready for commercial/non-commercial use.
Third-Party Community Consideration:
This model is not owned or developed by NVIDIA. This model has been developed and built to a third-party’s requirements for this application and use case. See link to Non-NVIDIA Qwen/Qwen2.5-Coder-32B-Instruct
License/Terms of Use
GOVERNING TERMS: The NIM container is governed by the NVIDIA Software License Agreement and the Product-Specific Terms for NVIDIA AI Products. The model is governed by the NVIDIA Community Model License Agreement.
ADDITIONAL INFORMATION: Apache License Version 2.0
Deployment Geography
Global
Release Date
Huggingface [06/09/2025] via https://huggingface.co/Qwen/Qwen2.5-Coder-32B-Instruct
Qwen2.5-Coder
Qwen2.5-Coder Container includes the following model:
Model Name & Link: Qwen2.5-Coder
Use Case: This is the core instruction-tuned model, designed for conversational code generation, reasoning, and fixing bugs based on natural language prompts.
How to Pull the Model: Automatic
##Deployment Details:
Visit the NIM Container LLM page for release documentation, deployment guides, and more.
Our AI models are designed and/or optimized to run on NVIDIA GPU-accelerated systems. By leveraging NVIDIA’s hardware (e.g. GPU cores) and software frameworks (e.g., CUDA libraries), the model achieves faster training and inference times compared to CPU-only solutions.
Reference(s):
N/A
Container Version(s):
nvcr.io/nvstaging/nim/qwen2.5-coder-32b-instruct:1.8.5-30597964
Ethical Considerations:
NVIDIA believes Trustworthy AI is a shared responsibility and we have established policies and practices to enable development for a wide array of AI applications. When downloaded or used in accordance with our terms of service, developers should work with their internal developer team to ensure this model meets requirements for the relevant industry and use case and addresses unforeseen product misuse. Please report security vulnerabilities or NVIDIA AI Concerns here.
Get Help
NVIDIA Developer Community Forum
For support, Visit the NVIDIA Developer Community Forum
End of Support — "This artifact is no longer supported. NVIDIA strongly recommends artifacts that are up to date and supported"