Resource
This notebook demonstrates how to optimize a fine-tuned BERT TF checkpoint to TensorRT and then how to deploy it for inference using Triton inference server on OpenShift cluster.
Use the NGC CLI to download:
Copied!
1 Version
This notebook demonstrates how to optimize a fine-tuned BERT TF checkpoint to TensorRT and then how to deploy it for inference using Triton inference server on OpenShift cluster.