NVIDIA
NVIDIA
Text Classification Notebook
Resource
NVIDIA
NVIDIA
Text Classification Notebook

End to End sample workflow for Text Classification starting with training in TLT and deployment using Jarvis.

text-classification-deployment.ipynb

Deploying Text Classification Model in Jarvis

Transfer Learning Toolkit (TLT) provides the capability to export your model in a format that can deployed using NVIDIA Jarvis, a highly performant application framework for multi-modal conversational AI services using GPUs.

This tutorial explores taking an .ejrvs model, the result of tlt text_classification export command, and leveraging the Jarvis ServiceMaker framework to aggregate all the necessary artifacts for Jarvis deployment to a target environment. Once the model is deployed in Jarvis, you can issue inference requests to the server. We will demonstrate how quick and straightforward this whole process is.

Learning Objectives

In this notebook, you will learn how to:

  • Use Jarvis ServiceMaker to take a TLT exported .ejrvs and convert it to .jmir
  • Deploy the model(s) locally on the Jarvis Server
  • Send inference requests from a demo client using Jarvis API bindings..

Pre-requisites

To follow along, please make sure:

  • You have access to NVIDIA NGC, and are able to download the Jarvis Quickstart resources
  • Have an .ejrvs model file that you wish to deploy. You can obtain this from tlt <task> export (with export_format=JARVIS). Please refer the tutorial on Text Classification using Transfer Learning Toolkit for more details on training and exporting an .ejrvs model.

Jarvis ServiceMaker

Servicemaker is the set of tools that aggregates all the necessary artifacts (models, files, configurations, and user settings) for Jarvis deployment to a target environment. It has two main components as shown below:

1. Jarvis-build

This step helps build a Jarvis-ready version of the model. It’s only output is an intermediate format (called a JMIR) of an end to end pipeline for the supported services within Jarvis. We are taking a ASR QuartzNet Model in consideration

jarvis-build is responsible for the combination of one or more exported models (.ejrvs files) into a single file containing an intermediate format called Jarvis Model Intermediate Representation (.jmir). This file contains a deployment-agnostic specification of the whole end-to-end pipeline along with all the assets required for the final deployment and inference. Please checkout the documentation to find out more.

In [ ]:
# IMPORTANT: UPDATE THESE PATHS 

# ServiceMaker Docker
JARVIS_SM_CONTAINER = "<add container name>"

# Directory where the .ejrvs model is stored $MODEL_LOC/*.ejrvs
MODEL_LOC = "<add path to model location>"

# Name of the .erjvs file
MODEL_NAME = "<add model name>"

# Key that model is encrypted with, while exporting with TLT
KEY = "<add encryption key used for trained model>"
In [ ]:
# Get the ServiceMaker docker
! docker pull $JARVIS_SM_CONTAINER
In [ ]:
# Syntax: jarvis-build <task-name> output-dir-for-jmir/model.jmir:key dir-for-ejrvs/model.ejrvs:key
# jarvis-build text_classification \
#            --domain_name="<your custom domain name>" \
#            /servicemaker-dev/<jmir_filename>:<encryption_key> \
#            /servicemaker-dev/<ejrvs_filename>:<encryption_key>

! docker run --rm --gpus 0 -v $MODEL_LOC:/data $JARVIS_SM_CONTAINER -- \
            jarvis-build text_classification -f /data/tc-model.jmir:$KEY /data/$MODEL_NAME:$KEY

NOTE: Above, tc-model.ejrvs is the text classification model obtained from tlt text_classification export.

2. Jarvis-deploy

The deployment tool takes as input one or more Jarvis Model Intermediate Representation (JMIR) files and a target model repository directory. It creates an ensemble configuration specifying the pipeline for the execution and finally writes all those assets to the output model repository directory.

In [ ]:
# Syntax: jarvis-deploy -f dir-for-jmir/model.jmir:key output-dir-for-repository
! docker run --rm --gpus 0 -v $MODEL_LOC:/data $JARVIS_SM_CONTAINER -- \
            jarvis-deploy -f /data/tc-model.jmir:$KEY /data/models

Start Jarvis Server

Once the model repository is generated, we are ready to start the Jarvis server. From this step onwards you need to download the Jarvis QuickStart Resource from NGC. Set the path to the directory here:

In [ ]:
# Set the Jarvis QuickStart directory
JARVIS_DIR = "<Path to the uncompressed folder downloaded from quickstart(include the folder name)>"

Next, we modify config.sh to enable relevant Jarvis services (asr for QuartzNet Model), provide the encryption key, and path to the model repository (jarvis_model_loc) generated in the previous step among other configurations.

For instance, if above the model repository is generated at $MODEL_LOC/models, then you can specify jarvis_model_loc as the same directory as MODEL_LOC

Pretrained versions of models specified in models_asr/nlp/tts are fetched from NGC. Since we are using our custom model, we can comment it in models_asr (and any others that are not relevant to your use case).

config.sh snippet

# Enable or Disable Jarvis Services 
service_enabled_asr=false                                                      ## MAKE CHANGES HERE
service_enabled_nlp=true                                                      ## MAKE CHANGES HERE
service_enabled_tts=false                                                     ## MAKE CHANGES HERE

# Specify one or more GPUs to use
# specifying more than one GPU is currently an experimental feature, and may result in undefined behaviours.
gpus_to_use="device=0"

# Specify the encryption key to use to deploy models
MODEL_DEPLOY_KEY="tlt_encode"                                                  ## MAKE CHANGES HERE

# Locations to use for storing models artifacts
#
# If an absolute path is specified, the data will be written to that location
# Otherwise, a docker volume will be used (default).
#
# jarvis_init.sh will create a `jmir` and `models` directory in the volume or
# path specified. 
#
# JMIR ($jarvis_model_loc/jmir)
# Jarvis uses an intermediate representation (JMIR) for models
# that are ready to deploy but not yet fully optimized for deployment. Pretrained
# versions can be obtained from NGC (by specifying NGC models below) and will be
# downloaded to $jarvis_model_loc/jmir by `jarvis_init.sh`
# 
# Custom models produced by NeMo or TLT and prepared using jarvis-build
# may also be copied manually to this location $(jarvis_model_loc/jmir).
#
# Models ($jarvis_model_loc/models)
# During the jarvis_init process, the JMIR files in $jarvis_model_loc/jmir
# are inspected and optimized for deployment. The optimized versions are
# stored in $jarvis_model_loc/models. The jarvis server exclusively uses these
# optimized versions.
jarvis_model_loc="<add path>"                              ## MAKE CHANGES HERE (Replace with MODEL_LOC)                      
In [ ]:
# Ensure you have permission to execute these scripts.
! cd $JARVIS_DIR && chmod +x ./jarvis_init.sh && chmod +x ./jarvis_start.sh
In [ ]:
# Run Jarvis Init. This will fetch the containers/models
# YOU CAN SKIP THIS STEP IF YOU DID JARVIS DEPLOY
! cd $JARVIS_DIR && ./jarvis_init.sh config.sh
In [ ]:
# Run Jarvis Start. This will deploy the model(s).
! cd $JARVIS_DIR && bash jarvis_start.sh config.sh

Run Inference

Once the Jarvis server is up and running with your models, you can send inference requests querying the server.

To send GRPC requests, you can install Jarvis Python API bindings for client. This is available as a pip .whl with the QuickStart.

In [ ]:
# IMPORTANT: Set the name of the whl file
JARVIS_API_WHL = "<add jarvis api .whl file name>"
In [ ]:
# Install client API bindings
!cd $JARVIS_DIR && pip install $JARVIS_API_WHL

The following code sample shows how you can perform inference using Jarvis Python API gRPC bindings:

In [ ]:
import grpc
import argparse
import os
import jarvis_api.jarvis_nlp_core_pb2 as jcnlp
import jarvis_api.jarvis_nlp_core_pb2_grpc as jcnlp_srv
import jarvis_api.jarvis_nlp_pb2 as jnlp
import jarvis_api.jarvis_nlp_pb2_grpc as jnlp_srv


class BertTextClassifyClient(object):
    def __init__(self, grpc_server, model_name):
        # generate the correct model based on precision and whether or not ensemble is used
        print("Using model: {}".format(model_name))

        self.model_name = model_name
        self.channel = grpc.insecure_channel(grpc_server)
        self.jarvis_nlp = jcnlp_srv.JarvisCoreNLPStub(self.channel)

        self.has_bos_eos = False

    # use the text_classification network to return top-1 classes for intents/sequences
    def postprocess_labels_server(self, ct_response):
        results = []

        for i in range(0, len(ct_response.results)):
            intent_str = ct_response.results[i].labels[0].class_name
            intent_conf = ct_response.results[i].labels[0].score

            results.append((intent_str, intent_conf))

        return results

    # accept a list of strings, return a list of tuples ('intent', scores)
    def run(self, input_strings):
        if isinstance(input_strings, str):
            # user probably passed a single string instead of a list/iterable
            input_strings = [input_strings]

        # get intent of the query
        request = jcnlp.TextClassRequest()
        request.model.model_name = self.model_name
        for q in input_strings:
            request.text.append(q)
        ct_response = self.jarvis_nlp.ClassifyText(request)

        return self.postprocess_labels_server(ct_response)


def run_text_classify(server, model, query):
    print("Client app to test text classification on Jarvis")
    client = BertTextClassifyClient(server, model_name=model)
    result = client.run(query)
    print(result)
In [ ]:
# Model Name will depend on the dataset and the domain on which the model was trained. 
# Please check `docker logs <container name>` and replace is accordingly (There will 
# be a table of models with their status displayed next to them) Check the documentation
# for more information.

run_text_classify(server="localhost:50051",
                model="<Enter Model Name>",
                query="How is the weather tomorrow?")

NOTE: You could also run the above inference code from inside the Jarvis Client container. The QuickStart provides a script jarvis_start_client.sh to run the container. It has more examples for different services.

You can stop all docker container before shutting down the jupyter kernel. Caution: The following command will stop all running containers

In [ ]:
! docker stop $(docker ps -a -q)

What's next?

You could train your own custom models in TLT and deploy them in Jarvis! You could scale up your deployment using Kubernetes with the Jarvis AI Services Helm Chart, which will pull the relevant Images and download model artifacts from NGC, generate the model repository, start and expose the Jarvis speech services.

NVIDIA uses cookies to improve your experience on our web site. We and our third-party partners also use cookies and other tools to collect and record information you provide as well as information about your interactions with our websites for performance improvement, analytics, and to assist in marketing efforts. By clicking "Accept All", you consent to our use of cookies and other tools as described in our Cookie Policy. You can manage your cookie settings by clicking on "Manage Settings." By continuing to use this site or by clicking one of the buttons below, you agree to our Terms of Service (which contains important waivers). Please see our Privacy Policy for more information on our privacy practices.