NVIDIA inference platform

E892043

The NVIDIA inference platform is a comprehensive suite of hardware and software tools designed to accelerate and optimize AI model deployment and real-time inference across data center, edge, and embedded environments.

All labels observed (3)

How this entity was disambiguated

Statements (69)

Predicate Object
instanceOf AI inference platform ⓘ
software and hardware platform ⓘ
developer NVIDIA ⓘ
linked to: NVIDIA Corporation
includesComponent NVIDIA AI Enterprise ⓘ
NVIDIA AI Workbench integration ⓘ
NVIDIA Base Command Manager ⓘ
NVIDIA BlueField DPUs ⓘ
NVIDIA CUDA ⓘ
NVIDIA DGX systems ⓘ
linked to: NVIDIA DGX

NVIDIA EGX platform ⓘ
NVIDIA GPU operator ⓘ
NVIDIA GPUs ⓘ
linked to: NVIDIA GPU hardware

NVIDIA Jetson platform ⓘ
NVIDIA NIM microservices ⓘ
linked to: NVIDIA NGC

NVIDIA NeMo microservices ⓘ
linked to: NVIDIA NeMo

NVIDIA TensorRT ⓘ
NVIDIA TensorRT-LLM ⓘ
linked to: NVIDIA TensorRT

NVIDIA Triton Inference Server ⓘ
NVIDIA cuDNN ⓘ
linked to: cuDNN

NVIDIA networking ⓘ
optimizationFeature FP16 mixed precision ⓘ
INT8 quantization ⓘ
dynamic batching ⓘ
layer fusion ⓘ
model ensemble execution ⓘ
precision calibration ⓘ
providesCapability autoscaling of inference workloads ⓘ
model optimization ⓘ
model serving ⓘ
multi-GPU inference ⓘ
multi-node inference ⓘ
observability and metrics for inference ⓘ
purpose accelerate AI inference ⓘ
optimize AI model deployment ⓘ
relatedTo NVIDIA AI platform ⓘ
NVIDIA training platform ⓘ
softwareStack CUDA ⓘ
linked to: NVIDIA CUDA

NVIDIA AI Enterprise ⓘ
TensorRT ⓘ
linked to: NVIDIA TensorRT

Triton Inference Server ⓘ
cuDNN ⓘ
supportsDeployment Kubernetes ⓘ
bare-metal servers ⓘ
cloud environments ⓘ
edge devices ⓘ
embedded modules ⓘ
on-premises data centers ⓘ
virtual machines ⓘ
supportsEnvironment data center ⓘ
edge ⓘ
embedded systems ⓘ
supportsFramework ONNX Runtime ⓘ
PyTorch ⓘ
TensorFlow ⓘ
XGBoost ⓘ
supportsModelFormat ONNX ⓘ
TensorFlow SavedModel ⓘ
TensorRT engine ⓘ
linked to: NVIDIA TensorRT

TorchScript ⓘ
linked to: PyTorch
supportsUseCase batch inference ⓘ
computer vision inference ⓘ
large language model inference ⓘ
online prediction services ⓘ
real-time inference ⓘ
recommender systems ⓘ
speech AI inference ⓘ
targetUser AI developers ⓘ
IT operations teams ⓘ
MLOps engineers ⓘ

How these facts were elicited

Referenced by (3)

Full triples — surface form annotated when it differs from this entity's canonical label.

NVIDIA TensorRT → componentOf → NVIDIA inference platform ⓘ
DeepStream SDK → optimizedFor → NVIDIA AI inference ⓘ
linked to: NVIDIA inference platform
NVIDIA inference platform → includesComponent → NVIDIA EGX platform ⓘ
linked to: NVIDIA inference platform