Run demanding AI workloads on enterprise-grade NVIDIA GPUs without purchasing and maintaining your own hardware.
Deploy your workloads in Malaysia when you require local data residency, lower regional latency and greater control over where your datasets, applications and model weights are stored and processed.
From cost-efficient inference on NVIDIA L20 GPUs to large-scale AI reasoning and training on NVIDIA H200, B300 and GB300 NVL72 infrastructure, our GPU as a Service offering supports organisations at every stage of their AI journey.
Avoid large upfront GPU purchases, hardware depreciation and the risk of investing in infrastructure that may become outdated.
Upgrade to newer GPU models or higher service tiers as your workload requirements grow, without replacing physical infrastructure.
Increase GPU capacity for model training, fine-tuning or high-volume inference, then scale down to smaller instances for development, testing and steady-state production workloads.
Access enterprise GPU infrastructure without managing procurement, installation, rack space, power, cooling, networking or hardware maintenance.
Keep AI workloads closer to Malaysia-based applications, users and data sources to support data-residency requirements and reduce regional network latency.
Use the same GPU infrastructure for model development, fine-tuning, evaluation, batch processing and production inference.
Provide a shared GPU resource pool for students, lecturers and researchers.
GPU resources can be allocated across classes, laboratories, research projects, hackathons and student experimentation, with user quotas, access controls and cost monitoring.
Create a centralised GPU platform that enables different departments to experiment with and deploy AI applications while maintaining governance, security and resource controls.
Conduct rapid prototyping, model benchmarking, model evaluation, internal tool development and innovation pilots without committing to long-term hardware ownership.
Fine-tune large language models using organisational or industry-specific knowledge, including financial, legal, healthcare, manufacturing and government datasets.
Train or optimise computer vision models for e-KYC, OCR, facial verification, liveness detection, quality inspection and document fraud detection.
Deploy scalable inference endpoints for:
These options are suitable for organisations that prioritise Malaysia-based deployment, local data residency and low-latency connectivity to Malaysia-based applications and users.
Final availability, instance configuration and commercial terms are subject to capacity confirmation.
A cost-efficient option for production inference, computer vision, smaller language models and moderate fine-tuning workloads.
Instance: ecs.gni3cl.5xlarge
Compute: 22 vCPU
System memory: 120GB RAM
GPU memory: 48GB per GPU
Pricing: Contact us for the best available price
Instance: ecs.gni3cl.11xlarge
Compute: 44 vCPU
System memory: 240GB RAM
GPU memory: 48GB per GPU
Total GPU memory: 96GB
Pricing: Contact us for the best available price
Instance: ecs.gni3cl.22xlarge
Compute: 90 vCPU
System memory: 480GB RAM
GPU memory: 48GB per GPU
Total GPU memory: 192GB
Pricing: Contact us for the best available price
Instance: ecs.gni3cl.45xlarge
Compute: 180 vCPU
System memory: 960GB RAM
GPU memory: 48GB per GPU
Total GPU memory: 384GB
Pricing: Contact us for the best available price
A high-performance Hopper-generation GPU for demanding AI training, large-scale fine-tuning and production inference.
Instance: ecs.hpcpni3h.42xlarge
Compute: 168 vCPU
System memory: 1,960GB RAM
GPU memory: 80GB per GPU
Total GPU memory: 640GB
Local storage: 3.84TB × 8
Network: 400G × 8 InfiniBand
Pricing: Contact us for the best available price
The NVIDIA H100 SXM configuration provides 80GB of HBM3 memory per GPU and supports high-speed NVLink connectivity for multi-GPU workloads. (NVIDIA)
Designed for memory-intensive large language models, RAG, larger context windows, high-throughput inference and advanced model fine-tuning.
GPU: NVIDIA H200 Tensor Core GPU
GPU memory: 141GB HBM3e per GPU
Memory bandwidth: 4.8TB/s per GPU
Recommended deployment: Dedicated multi-GPU node
CPU, RAM, storage and network: Configured according to workload requirements
Pricing: Contact us for the best available price
An eight-GPU H200 deployment provides approximately 1.13TB of aggregate GPU memory, making it suitable for larger models, larger batch sizes and workloads that are constrained by GPU memory.
NVIDIA specifies 141GB of HBM3e memory and 4.8TB/s of memory bandwidth for each H200 GPU. (NVIDIA)
A premium GPU tier for large-scale AI reasoning, agentic AI, advanced post-training, multimodal models and high-throughput training and inference.
GPU: 8 × NVIDIA B300 Blackwell Ultra GPUs
GPU memory: 288GB HBM3e per GPU
Total GPU memory: Approximately 2.3TB
FP8 training performance: Up to 72 PFLOPS per eight-GPU system
FP4 inference performance: Up to 144 PFLOPS per eight-GPU system
Interconnect: Fifth-generation NVIDIA NVLink and NVSwitch
Network: Up to 800Gb/s InfiniBand or Ethernet connectivity, depending on configuration
CPU, RAM and storage: Configured according to deployment requirements
Pricing: Contact us for the best available price
The NVIDIA DGX B300 reference system combines eight B300 GPUs, each with 288GB of HBM3e memory, for approximately 2.3TB of total GPU memory. (NVIDIA Docs)
The highest infrastructure tier for frontier-scale AI reasoning, large multimodal models, distributed training and extremely high-throughput inference.
Unlike a conventional GPU virtual machine, the GB300 NVL72 is a fully liquid-cooled, rack-scale system that operates as one large NVLink compute domain.
GPU: 72 × NVIDIA Blackwell Ultra GPUs
CPU: 36 × NVIDIA Grace CPUs
GPU memory: Approximately 20TB
GPU memory bandwidth: Up to 576TB/s aggregate
CPU memory: Approximately 17TB LPDDR5X
Total fast memory: Approximately 37TB
CPU cores: 2,592 Arm Neoverse V2 cores
NVLink bandwidth: Approximately 130TB/s
Cooling: Fully liquid-cooled rack-scale architecture
Deployment: Dedicated rack or reserved infrastructure capacity
Pricing: Contact us for availability and customised commercial terms
GB300 NVL72 is designed for AI reasoning, test-time scaling, agentic AI, trillion-parameter-class models and large-scale AI factory deployments.
NVIDIA specifies that the GB300 NVL72 integrates 72 Blackwell Ultra GPUs and 36 Grace CPUs, with approximately 20TB of GPU memory, 576TB/s aggregate GPU memory bandwidth and 130TB/s of NVLink bandwidth. (NVIDIA)
Best for:
Why choose it:
L20 provides a practical balance between GPU memory, performance and cost for organisations deploying mainstream enterprise AI applications.
Best for:
Why choose it:
H100 remains a widely supported enterprise GPU platform with a mature software ecosystem and strong training and inference performance.
Best for:
Why choose it:
The H200 increases GPU memory to 141GB and provides 4.8TB/s of memory bandwidth, reducing memory bottlenecks for large models and data-intensive workloads. (NVIDIA Newsroom)
Best for:
Why choose it:
Each B300 GPU provides 288GB of HBM3e memory. An eight-GPU system delivers approximately 2.3TB of total GPU memory, providing substantially more room for large models, context windows and batch sizes. (NVIDIA Docs)
Best for:
Why choose it:
GB300 NVL72 connects 72 Blackwell Ultra GPUs through a rack-scale NVLink architecture, providing approximately 20TB of GPU memory and 130TB/s of NVLink bandwidth. (NVIDIA)
Choose L20 when cost-efficient inference and moderate fine-tuning are your main priorities.
Choose H100 when you need proven high-performance training and enterprise-grade multi-GPU computing.
Choose H200 when your workload requires more GPU memory for larger models, longer context windows or larger batch sizes.
Choose B300 when you are developing advanced reasoning models, agentic AI systems, multimodal applications or high-throughput generative AI services.
Choose GB300 NVL72 when you require dedicated rack-scale AI infrastructure for frontier models, large-scale reasoning or AI factory deployments.
Speak with our team to identify the most suitable GPU configuration for your AI workload.
Email: [email protected]
无需购买及维护自有硬件,即可使用企业级 NVIDIA GPU 运行高性能 AI 工作负载。
当企业需要将数据部署在马来西亚、降低区域网络延迟,以及更严格地控制数据集、应用程序及模型权重的存储与处理位置时,可选择我们的马来西亚 GPU 基础设施。
从具备成本效益的 NVIDIA L20 推理服务,到适合大型 AI 推理与训练的 NVIDIA H200、B300 及 GB300 NVL72,我们的 GPU 即服务方案可支持企业不同阶段的 AI 发展需求。
无需承担高额的 GPU 前期采购成本、硬件折旧,以及设备快速过时所带来的投资风险。
当 AI 工作负载增加或新一代 GPU 推出时,可直接升级 GPU 型号或服务等级,无需更换自有实体基础设施。
在模型训练、微调或高流量推理期间增加 GPU 容量;项目完成后,则可缩减至较小的开发、测试或稳定生产环境。
无需自行处理采购、设备安装、机架空间、供电、冷却、网络配置及硬件维护,即可快速取得企业级 GPU 资源。
让 AI 工作负载更接近位于马来西亚的应用程序、用户及数据源,以支持数据驻留需求并降低区域网络延迟。
同一套 GPU 基础设施可用于模型开发、微调、评估、批量处理及生产环境推理。
为学生、讲师及研究人员提供共享 GPU 资源池。
GPU 资源可分配给课程、实验室、研究项目、黑客松及学生 AI 实验,并支持用户配额、访问权限及成本监控。
建立中央 GPU 平台,让不同部门在统一的安全、治理及资源管理机制下进行 AI 实验、开发及部署。
无需长期持有硬件,即可进行快速原型开发、模型性能测试、模型评估、内部工具开发及创新试点。
使用企业或行业数据微调大语言模型,包括金融、法律、医疗、制造及政府相关知识。
训练及优化计算机视觉模型,以支持电子身份认证、OCR、脸部验证、活体检测、品质检查及文件欺诈识别。
部署可扩展的推理端点,以支持:
以下选项适合重视马来西亚本地部署、数据驻留,以及与马来西亚应用程序和用户之间低延迟连接的机构。
最终供应情况、实例配置及商业条款须视实际容量而定。
适合生产环境推理、计算机视觉、较小型语言模型及中等规模模型微调的高性价比选择。
实例规格: ecs.gni3cl.5xlarge
计算资源: 22 vCPU
系统内存: 120GB RAM
GPU 显存: 每张 GPU 48GB
价格: 请联系我们以获取最优惠价格
实例规格: ecs.gni3cl.11xlarge
计算资源: 44 vCPU
系统内存: 240GB RAM
GPU 显存: 每张 GPU 48GB
GPU 总显存: 96GB
价格: 请联系我们以获取最优惠价格
实例规格: ecs.gni3cl.22xlarge
计算资源: 90 vCPU
系统内存: 480GB RAM
GPU 显存: 每张 GPU 48GB
GPU 总显存: 192GB
价格: 请联系我们以获取最优惠价格
实例规格: ecs.gni3cl.45xlarge
计算资源: 180 vCPU
系统内存: 960GB RAM
GPU 显存: 每张 GPU 48GB
GPU 总显存: 384GB
价格: 请联系我们以获取最优惠价格
适合高强度 AI 训练、大规模模型微调及生产环境推理的高性能 Hopper 架构 GPU。
实例规格: ecs.hpcpni3h.42xlarge
计算资源: 168 vCPU
系统内存: 1,960GB RAM
GPU 显存: 每张 GPU 80GB
GPU 总显存: 640GB
本地存储: 3.84TB × 8
网络: 400G × 8 InfiniBand
价格: 请联系我们以获取最优惠价格
NVIDIA H100 SXM 配置为每张 GPU 提供 80GB HBM3 显存,并支持高速 NVLink 多 GPU 连接。(NVIDIA)
适合对显存要求较高的大语言模型、RAG、长上下文窗口、高吞吐量推理及高级模型微调。
GPU: NVIDIA H200 Tensor Core GPU
GPU 显存: 每张 GPU 141GB HBM3e
显存带宽: 每张 GPU 4.8TB/s
建议部署方式: 专属多 GPU 节点
CPU、RAM、存储及网络: 根据工作负载需求配置
价格: 请联系我们以获取最优惠价格
8 张 H200 GPU 可提供约 1.13TB GPU 总显存,适合运行更大型的模型、更大的批处理规模,以及受 GPU 显存限制的工作负载。
根据 NVIDIA 的官方规格,每张 H200 GPU 配备 141GB HBM3e 显存及 4.8TB/s 显存带宽。(NVIDIA)
适合大规模 AI 推理、智能体 AI、高级模型后训练、多模态模型,以及高吞吐量训练和推理的高端 GPU 服务等级。
GPU: 8 张 NVIDIA B300 Blackwell Ultra GPU
GPU 显存: 每张 GPU 288GB HBM3e
GPU 总显存: 约 2.3TB
FP8 训练性能: 8 GPU 系统最高可达 72 PFLOPS
FP4 推理性能: 8 GPU 系统最高可达 144 PFLOPS
互连技术: 第五代 NVIDIA NVLink 及 NVSwitch
网络: 根据配置最高可支持 800Gb/s InfiniBand 或以太网连接
CPU、RAM 及存储: 根据部署需求配置
价格: 请联系我们以获取最优惠价格
NVIDIA DGX B300 参考系统配备 8 张 B300 GPU,每张 GPU 配备 288GB HBM3e 显存,总 GPU 显存约为 2.3TB。(NVIDIA Docs)
适合前沿级 AI 推理、大型多模态模型、分布式训练及超高吞吐量推理的最高端基础设施等级。
GB300 NVL72 并非一般的 GPU 虚拟机实例,而是一套采用全液冷设计的机架级系统,并通过 NVLink 连接为一个大型计算域。
GPU: 72 张 NVIDIA Blackwell Ultra GPU
CPU: 36 个 NVIDIA Grace CPU
GPU 显存: 约 20TB
GPU 显存总带宽: 最高约 576TB/s
CPU 内存: 约 17TB LPDDR5X
高速内存总量: 约 37TB
CPU 核心: 2,592 个 Arm Neoverse V2 核心
NVLink 带宽: 约 130TB/s
冷却系统: 全液冷机架级架构
部署方式: 专属机架或预留基础设施容量
价格: 请联系我们确认供应情况并获取定制商业方案
GB300 NVL72 专为 AI 推理、测试时扩展、智能体 AI、万亿参数级模型及大型 AI 工厂部署而设计。
根据 NVIDIA 的官方规格,GB300 NVL72 整合了 72 张 Blackwell Ultra GPU 及 36 个 Grace CPU,并提供约 20TB GPU 显存、最高 576TB/s GPU 显存总带宽及 130TB/s NVLink 带宽。(NVIDIA)
最适合:
选择原因:
L20 在 GPU 显存、性能及成本之间取得良好平衡,适合部署主流企业 AI 应用。
最适合:
选择原因:
H100 拥有成熟的软件生态系统,并具备强大的训练及推理性能,仍是广泛使用的企业级 GPU 平台。
最适合:
选择原因:
H200 将每张 GPU 的显存提升至 141GB,并提供 4.8TB/s 显存带宽,有助于减少大型模型及数据密集型工作负载的显存瓶颈。(NVIDIA Newsroom)
最适合:
选择原因:
每张 B300 GPU 配备 288GB HBM3e 显存。8 GPU 系统可提供约 2.3TB GPU 总显存,为大型模型、长上下文窗口及更大的批处理规模提供充足空间。(NVIDIA Docs)
最适合:
选择原因:
GB300 NVL72 通过机架级 NVLink 架构连接 72 张 Blackwell Ultra GPU,提供约 20TB GPU 显存及 130TB/s NVLink 带宽。(NVIDIA)
如主要需求是具备成本效益的推理及中等规模模型微调,可选择 L20。
如需要成熟的高性能模型训练及企业级多 GPU 计算,可选择 H100。
如需要更多 GPU 显存来运行大型模型、长上下文窗口或更大的批处理规模,可选择 H200。
如正在开发高级推理模型、智能体 AI、多模态应用或高吞吐量生成式 AI 服务,可选择 B300。
如需要专属的机架级 AI 基础设施,以支持前沿模型、大规模推理或 AI 工厂部署,可选择 GB300 NVL72。
欢迎与我们的团队联系,让我们根据您的 AI 工作负载推荐最合适的 GPU 配置。
电子邮件:[email protected]
微信:tanaikkeong
