TABLE OF CONTENTS
NVIDIA H100 SXM On-Demand
Welcome to Hyperstack Monthly Update
July was all about making Hyperstack easier to build with, easier to integrate into AI workflows and more reliable in production. From bringing our documentation directly into AI assistants to expanding AI Studio with multimodal capabilities and long-context models, every update helps you spend less time managing infrastructure and more time building.
New on Hyperstack
Check out what's new on Hyperstack:
Hyperstack Docs MCP Server
We introduced the Hyperstack Docs MCP Server, a read-only Model Context Protocol (MCP) server that connects MCP-compatible AI clients directly to Hyperstack documentation.
Instead of searching documentation manually, your AI assistant can retrieve answers from the official Hyperstack docs and include citation links back to the source pages. Hyperstack hosts the server, requires no authentication and is ready to use with compatible MCP clients.
New on AI Studio
Check out what's new on AI Studio this month:
Kimi K3 as a Third-Party Hosted Model
Kimi K3 is now available as a third-party model on AI Studio. You can launch it directly from the Text Playground or access it through the API using moonshotai/Kimi-K3.
The model accepts both text and image inputs, generates text responses and supports a context window of up to 1 million tokens. It is ideal for applications involving lengthy documents, large codebases and long-running conversations.
Image-to-Text Support
AI Studio now supports image-to-text workflows across both the Playground and the API. You can attach images in the Text Playground and use supported models to describe visual content, answer questions about images, extract information from screenshots or analyse documents.
The Model Catalog has also been expanded with dedicated image-to-text models, making it easier to discover models that support multimodal inputs.
Conversation History for the Text Playground
Text Playground conversations are now automatically saved. You can browse previous sessions, rename conversations, delete those you no longer need and revisit the complete message history, including any images shared during the conversation.
Conversation management is also available through the new Conversations APIs, making it easier to build applications that store and manage chat history programmatically.
Explore Image Playground →
Latest Improvements
July also included several reliability improvements across Hyperstack:
-
Improved cluster reconciliation: Resolved an issue where failed scale-up operations could leave Kubernetes clusters stuck in the RECONCILING state, preventing further cluster operations.
-
Master node deletion fix: Resolved an issue where removing the first master node could leave remaining nodes pointing to an invalid control plane address, potentially making the cluster unreachable.
-
Node group cleanup: Improved node group deletion to prevent orphaned nodes from remaining after a node group had been removed.
-
API key limit enforcement: Improved API key creation so account limits are enforced consistently, even when multiple API keys are created simultaneously.
New on our Blog
Check out the latest tutorials and blogs on Hyperstack:
Deploy Kimi K3 on GPU Cloud for Multi-Node 2.8T Inference
A Comprehensive Guide
Kimi K3 is a 2.8 trillion parameter sparse Mixture-of-Experts model with a one million token context window, making it one of the largest open-weight models available. This guide shows how to deploy it across 32 NVIDIA H100 GPUs with vLLM on Hyperstack.
Deploy Hy3 on GPU Cloud for Multi-Node 295B Inference
A Comprehensive Guide
Tencent Hy3 is too large to fit on a single 8x NVIDIA H100 node, requiring a distributed 16-GPU deployment for full BF16 inference. This guide shows how to deploy it with vLLM on Hyperstack, from setup to a live endpoint.
The Race to Run AI at Scale Is Now a Race
For More Tokens Per Watt
For years, AI infrastructure was judged by GPU count. Today, the more important question is how many tokens your server produces per watt. As AI shifts from training to large-scale inference, efficiency, latency and cost per million tokens are becoming the metrics that determine profitability.
Help Shape the Future of Hyperstack
Great products are built with the people who use them. If there’s something you would like to see on Hyperstack, whether it is a new feature, workflow improvement or integration that would make your work easier, we would love to hear about it.
Your feedback helps us prioritise what matters most and build a platform that works better for the community.
That's it for this Monthly Update! Stay tuned for more updates and subscribe to our newsletter below for exclusive AI and GPU insights delivered to your inbox!
Subscribe to Hyperstack!
Enter your email to get updates to your inbox every week
Get Started
Ready to build the next big thing in AI?