NVIDIA Triton Inference Server

E234124

NVIDIA Triton Inference Server is an open-source, production-ready platform for serving and scaling AI model inference across GPUs and CPUs with support for multiple frameworks and deployment environments.

All labels observed (1)

Label Occurrences
NVIDIA Triton Inference Server canonical 3

How this entity was disambiguated

Statements (57)

Predicate Object
instanceOf AI inference server
model serving software
open-source software
developer NVIDIA
linked to: NVIDIA Corporation
hasFeature GPU-aware scheduling
auto-scaling integration
dynamic model loading
model ensemble support
model repository
model warmup
request batching
license BSD 3-Clause License
linked to: BSD license
officialWebsite https://developer.nvidia.com/nvidia-triton-inference-server
programmingLanguage C++
Python
repositoryUrl https://github.com/triton-inference-server/server
supportsBatching true
supportsCloudPlatform NVIDIA AI Enterprise
NVIDIA DGX systems
linked to: NVIDIA DGX

major public clouds
supportsConcurrentModelExecution true
supportsDeploymentEnvironment Docker
Kubernetes
bare metal
virtual machines
supportsDynamicBatching true
supportsFormat C++ model
ONNX
Python model
TensorFlow GraphDef
TensorFlow SavedModel
TensorRT engine
linked to: NVIDIA TensorRT

TorchScript
linked to: PyTorch
supportsFramework C++ backend
ONNX Runtime
OpenVINO
PyTorch
Python backend
TensorFlow
TensorRT
linked to: NVIDIA TensorRT
supportsGRPCProtocol true
supportsHardware CPU
GPU
supportsHTTPProtocol true
supportsMetrics CPU utilization
GPU utilization
Prometheus
request latency
throughput
supportsModelType NLP models
computer vision models
recommendation models
supportsModelVersioning true
supportsMultiModelServing true
useCase batch inference
online inference
real-time AI serving

How these facts were elicited

Referenced by (3)

Full triples — surface form annotated when it differs from this entity's canonical label.

NVIDIA DGX supports NVIDIA Triton Inference Server
NVIDIA AI Enterprise includes NVIDIA Triton Inference Server
NVIDIA TensorRT integratesWith NVIDIA Triton Inference Server