d-Matrix announced it will use NVIDIA NVLink Fusion to connect its next-generation Raptor XPUs to NVIDIA's AI infrastructure platform, including NVLink scale-up and Spectrum-X networking and the MGX rack architecture. The company said this gives it a path to deploy Raptor at large scale alongside NVIDIA systems.
AWS launched model caching for Amazon SageMaker Inference on HyperPod. The feature pre-loads model weights and container images onto cluster nodes so pods can read from local NVMe storage at about 7 GB/s instead of downloading over the network, letting pods typically start serving traffic in seconds rather than tens of minutes, AWS says.
Databricks unveiled Proteus, a system that uses agents to generate specialized GPU kernels and subjects candidates to controlled validation and repeated benchmarking. The company says individual kernels generated for Qwen 3.5 122B were 1.8 to 5.2 times faster than the best implementations available in vLLM.
Anthropic selected Google Cloud as its cloud provider under a partnership announced on Feb. 3, 2023. The companies plan to co-develop AI computing systems, while Anthropic says it will use Google Cloud’s GPU and TPU clusters to train, scale and deploy its AI systems, including Claude.
Only at high, steady utilisation. Self-hosting converts a variable per-token cost into a fixed hourly cost, so it wins when accelerators stay busy and loses badly when they idle. The honest comparison prices the full stack — accelerator hours, redundancy, engineering time, and the evaluation work needed to confirm the smaller model is good enough — against the API bill for the same traffic. Sovereignty, data residency and latency floors are separate reasons that can justify self-hosting regardless of the arithmetic.
AI accelerators draw far more power per rack than traditional servers, and that power has to be delivered, cooled and paid for continuously. Training a large model is a one-off spike; serving it to millions of users is a permanent load, and inference is what dominates energy use over a deployed model's life.