Developed jointly by Quantiphi Inc and NVIDIA, this repository provides an example of an LLM-based RAG (Retriever Augmented Generation) pipeline to accelerate and simplify the implementation of chatbot solutions in the Telco industry.
The workflow was tested on a single NVIDIA A100 GPU (80GB GPU Memory)
- Langchain
- TensorRT-LLM
- Triton Inference Server
- Milvus
- Hugging Face
- Streamlit
- FastAPI
Cloning the repo containes services along with tensorrtllm_backend as a submodule which comes with submodules that can be updated with the commands below.
cd tensorrtllm_backend/
git lfs install
git lfs pull
git submodule update --init --recursive
cd ../Clone and place a embedding model of choice under backend/embedding_model. We have used BAAI/bge-base-en-v1.5.
any changes to the embedding model has to be updated in backend/config.yml file.
git clone https://huggingface.co/BAAI/bge-base-en-v1.5 backend/embedding_model/BAAI/bge-base-en-v1.5We have the fronend, backend and Milvus DB as a microservice which can be started using the compose file provided along with the project.
docker compose up -dStarting the backend would also mount './dataset' folder inside the container. To index the data into the vector database hit the /ingest_data endpoint with the dataset folder. This would chunk the data and index the same for RAG Pipeline.
curl -X POST \
'http://localhost:9999/ingest_data' \
--header 'Accept: */*' \
--header 'Content-Type: application/json' \
--data-raw '{
"path":"/dataset"
}'In the below examples we show the commands for running a gated LLAMA2-13b-chat model from Hugging Face. We would need the Hugging Face token to be passed in the env variable of the below command. We assume that we already have the engine file for the model that we are deploying. Follow this document to create the engine file to create your own engine file. We can also adjust the GPU that is going to be used for Triton Inference Server by changing the --gpus argument. In our example, we have specified the GPU device id as 0.
Running this Docker container will add the container to the same network as other microservices, so that they can communicate with each other.
# Update HF_TOKEN
docker run --rm -it --env HF_TOKEN=hf_* \
--name=triton_server --network rag_accelerator \
-p8000:8000 -p8001:8001 -p8002:8002 --shm-size=2g \
--ulimit memlock=-1 --ulimit stack=67108864 \
--gpus '"device=0"' -v ./tensorrtllm_backend:/tensorrtllm_backend \
nvcr.io/nvidia/tritonserver:23.10-trtllm-python-py3 bashThe following commands are to be executed inside the Docker container that we started in the above section. We have to adjust the commands based on the model that we are using and the location where we put the engine files. This example assumes that we have 2 folders
- tensorrtllm_backend/tensorrt_llm/examples/llama/Llama-2-13b-chat-hf
- tensorrtllm_backend/tensorrt_llm/examples/llama/Llama-2-13b-chat-hf-engine
First one for the hugging face model and the second one for storing the engine file created by TensorRT-LLM.
Login with HF_TOKEN inside the container using huggingface-cli login --token $HF_TOKEN
Install the requirements of llama model.
pip install -r /tensorrtllm_backend/tensorrt_llm/examples/llama/requirements.txt && pip install protobuf
cp -R /tensorrtllm_backend/all_models/inflight_batcher_llm /opt/tritonserver/.
sed -i 's#${tokenizer_dir}#meta-llama/Llama-2-13b-hf#' /opt/tritonserver/inflight_batcher_llm/preprocessing/config.pbtxt && \
sed -i 's#${tokenizer_type}#llama#' /opt/tritonserver/inflight_batcher_llm/preprocessing/config.pbtxt && \
sed -i 's#${tokenizer_dir}#meta-llama/Llama-2-13b-hf#' /opt/tritonserver/inflight_batcher_llm/postprocessing/config.pbtxt && \
sed -i 's#${tokenizer_type}#llama#' /opt/tritonserver/inflight_batcher_llm/postprocessing/config.pbtxt && \
sed -i 's#${decoupled_mode}#true#' /opt/tritonserver/inflight_batcher_llm/tensorrt_llm/config.pbtxt && \
sed -i 's#${engine_dir}#/tensorrtllm_backend/tensorrt_llm/examples/llama/Llama-2-13b-chat-hf-engine/1-gpu/#' /opt/tritonserver/inflight_batcher_llm/tensorrt_llm/config.pbtxt
Run the command to start the triton server:
tritonserver --model-repository=/opt/tritonserver/inflight_batcher_llm
Now we have all the services ready. The frontend will be accessible on http://localhost:8501 port.
Quantiphi is an award-winning AI-first digital engineering company driven by the desire to reimagine and realize transformational opportunities at the heart of the business. Since its inception in 2013, Quantiphi has solved the toughest and most complex business problems by combining deep industry experience, disciplined cloud, and data-engineering practices, and cutting-edge artificial intelligence research to achieve accelerated and quantifiable business results. Learn more at www.quantiphi.com.



