NVIDIA TensorRT

E256947

NVIDIA TensorRT is a high-performance deep learning inference optimizer and runtime library designed to accelerate AI models on NVIDIA GPUs in production environments.

All labels observed (4)

Label Occurrences
TensorRT 3
NVIDIA TensorRT canonical 1
NVIDIA TensorRT (indirectly via shared primitives) 1

How this entity was disambiguated

Statements (49)

Predicate Object
instanceOf deep learning inference optimizer
inference runtime library
componentOf NVIDIA AI software stack
NVIDIA inference platform
designedFor production AI deployment
developer NVIDIA
linked to: NVIDIA Corporation
distribution NVIDIA Developer website
NVIDIA GPU Cloud containers
goal maximize inference throughput
minimize inference latency
integratesWith CUDA
linked to: NVIDIA CUDA

NVIDIA DeepStream
linked to: DeepStream SDK

NVIDIA Triton Inference Server
ONNX Runtime
PyTorch via ONNX export
TensorFlow via TensorRT integration
cuDNN
license proprietary
optimizedFor NVIDIA GPUs
linked to: Nvidia Maxwell GPU
primaryUse deep learning inference acceleration
providesFeature CUDA integration
calibration for INT8 quantization
dynamic shapes support
dynamic tensor memory management
graph optimizations
kernel auto-tuning
layer fusion
multi-stream execution
plugin layer mechanism
supportsDeploymentEnvironment cloud environments
edge devices
on-premises data centers
supportsHardware NVIDIA data center GPUs
NVIDIA embedded GPUs
NVIDIA gaming GPUs
supportsLanguageBinding C++
Python
supportsModelFormat NVIDIA framework-specific formats
ONNX
supportsPrecision FP16
FP32
INT8
TF32
targetDomain computer vision inference
natural language processing inference
recommendation systems inference
speech and audio inference
usedFor batch inference
real-time inference

How these facts were elicited

Referenced by (6)

Full triples — surface form annotated when it differs from this entity's canonical label.

NVIDIA AI Enterprise includes NVIDIA TensorRT
NVIDIA Jetson embedded modules supports TensorRT
linked to: NVIDIA TensorRT
cuDNN integratesWith NVIDIA TensorRT (indirectly via shared primitives)
linked to: NVIDIA TensorRT
Tensor Cores exposedThrough TensorRT
linked to: NVIDIA TensorRT
NVIDIA Triton Inference Server supportsFramework TensorRT
linked to: NVIDIA TensorRT
NVIDIA Triton Inference Server supportsFormat TensorRT engine
linked to: NVIDIA TensorRT