BLOG

Engineering Notes.

Notes on large-scale training, GPU infrastructure, and what we're learning from engineering communities.

The Hidden Frictions of Renting GPUs

Clicking "Rent GPU" is the easy part. Reading 77 Reddit accounts of what actually happens next — setup, storage, availability, and recovery — tells a different story.

MTP in Practice: Benchmarking Gemma's Speculative Decoding on a Real GPU

We ran Gemma-4-12B's MTP drafter on a consumer RTX 3060 and measured, turn by turn, exactly how much faster it gets. Part 2 of 2: the implementation/benchmark edition.

MTP: The Low-Risk, High-Reward Bet Behind Faster Generation

Qwen and Gemma both predict several tokens ahead to speed up generation — but they've built completely different machinery to do it. Part 1 of 2: the concept edition.

Getting Started with AI Agent Development: A Cloud Architecture Guide

From GPU selection to infrastructure design — the fundamentals every developer needs to know to run autonomous AI agents on their own infrastructure.

Stop comparing GPU clouds only by $/hour

The cheapest GPU instance is not always the cheapest way to finish a workload.

From Vibe Coding to GPU Spec-Driven Development.

The development style AI-era engineers need to know: why Vibe Coding hits a wall, and how Spec-Driven Development (SDD) picks up where it leaves off.

What AI engineers actually care about when choosing a GPU cloud.

We reviewed 100 Reddit threads on GPU cloud decisions. Price was the primary pain point in 23—but in 74, cost was not the primary decision criterion.

Get early access to Compute Cluster

Priority allocation and founding-customer pricing for early signups.