NVIDIA TensorRT

E256947

NVIDIA TensorRT is a high-performance deep learning inference optimizer and runtime library designed to accelerate AI models on NVIDIA GPUs in production environments.

All labels observed (7)

How this entity was disambiguated

Statements (49)

Predicate Object
instanceOf deep learning inference optimizer ⓘ
inference runtime library ⓘ
componentOf NVIDIA AI software stack ⓘ
NVIDIA inference platform ⓘ
designedFor production AI deployment ⓘ
developer NVIDIA ⓘ
linked to: NVIDIA Corporation
distribution NVIDIA Developer website ⓘ
NVIDIA GPU Cloud containers ⓘ
goal maximize inference throughput ⓘ
minimize inference latency ⓘ
integratesWith CUDA ⓘ
linked to: NVIDIA CUDA

NVIDIA DeepStream ⓘ
linked to: DeepStream SDK

NVIDIA Triton Inference Server ⓘ
ONNX Runtime ⓘ
PyTorch via ONNX export ⓘ
TensorFlow via TensorRT integration ⓘ
cuDNN ⓘ
license proprietary ⓘ
optimizedFor NVIDIA GPUs ⓘ
linked to: Nvidia Maxwell GPU
primaryUse deep learning inference acceleration ⓘ
providesFeature CUDA integration ⓘ
calibration for INT8 quantization ⓘ
dynamic shapes support ⓘ
dynamic tensor memory management ⓘ
graph optimizations ⓘ
kernel auto-tuning ⓘ
layer fusion ⓘ
multi-stream execution ⓘ
plugin layer mechanism ⓘ
supportsDeploymentEnvironment cloud environments ⓘ
edge devices ⓘ
on-premises data centers ⓘ
supportsHardware NVIDIA data center GPUs ⓘ
NVIDIA embedded GPUs ⓘ
NVIDIA gaming GPUs ⓘ
supportsLanguageBinding C++ ⓘ
Python ⓘ
supportsModelFormat NVIDIA framework-specific formats ⓘ
ONNX ⓘ
supportsPrecision FP16 ⓘ
FP32 ⓘ
INT8 ⓘ
TF32 ⓘ
targetDomain computer vision inference ⓘ
natural language processing inference ⓘ
recommendation systems inference ⓘ
speech and audio inference ⓘ
usedFor batch inference ⓘ
real-time inference ⓘ

How these facts were elicited

Referenced by (23)

Full triples — surface form annotated when it differs from this entity's canonical label.

NVIDIA AI Enterprise → includes → NVIDIA TensorRT ⓘ
NVIDIA Jetson embedded modules → supports → TensorRT ⓘ
linked to: NVIDIA TensorRT
cuDNN → integratesWith → NVIDIA TensorRT (indirectly via shared primitives) ⓘ
linked to: NVIDIA TensorRT
Tensor Cores → exposedThrough → TensorRT ⓘ
linked to: NVIDIA TensorRT
NVIDIA Triton Inference Server → supportsFramework → TensorRT ⓘ
linked to: NVIDIA TensorRT
NVIDIA Triton Inference Server → supportsFormat → TensorRT engine ⓘ
linked to: NVIDIA TensorRT
ONNX → ecosystem → TensorRT (via parsers) ⓘ
linked to: NVIDIA TensorRT
NVIDIA JetPack SDK → includesComponent → NVIDIA TensorRT ⓘ
DeepStream SDK → usesFramework → TensorRT ⓘ
linked to: NVIDIA TensorRT
CUDA libraries → component → TensorRT (CUDA-accelerated inference library) ⓘ
linked to: NVIDIA TensorRT
NVIDIA Developer website → hostsSDK → TensorRT ⓘ
linked to: NVIDIA TensorRT
NVIDIA A100 → supportsTechnology → TensorRT ⓘ
linked to: NVIDIA TensorRT
NVIDIA A40 → supports → NVIDIA TensorRT ⓘ
NVIDIA A30 → supports → TensorRT ⓘ
linked to: NVIDIA TensorRT
NVIDIA L40S → supports → NVIDIA TensorRT ⓘ
ONNX Runtime → supportsExecutionProvider → TensorRT ⓘ
linked to: NVIDIA TensorRT
NVIDIA CUDA-X AI → hasComponent → TensorRT ⓘ
linked to: NVIDIA TensorRT
NVIDIA CUDA-X AI → hasComponent → NVIDIA TensorRT-LLM ⓘ
linked to: NVIDIA TensorRT
NVIDIA inference platform → includesComponent → NVIDIA TensorRT-LLM ⓘ
linked to: NVIDIA TensorRT
NVIDIA inference platform → softwareStack → TensorRT ⓘ
linked to: NVIDIA TensorRT
NVIDIA inference platform → supportsModelFormat → TensorRT engine ⓘ
linked to: NVIDIA TensorRT
NVIDIA NGC → includes → NVIDIA TensorRT ⓘ