AWS launched model caching for Amazon SageMaker Inference on HyperPod. The feature pre-loads model weights and container images onto cluster nodes so pods can read from local NVMe storage at about 7 GB/s instead of downloading over the network, letting pods typically start serving traffic in seconds rather than tens of minutes, AWS says.
Databricks unveiled Proteus, a system that uses agents to generate specialized GPU kernels and subjects candidates to controlled validation and repeated benchmarking. The company says individual kernels generated for Qwen 3.5 122B were 1.8 to 5.2 times faster than the best implementations available in vLLM.