128GB of Local VRAM vs The Cloud: Cost Optimisation, Zero Telemetry, and What’s Possible with Local Dev Stacks
Token counts, cloud subscriptions, and API rate limits are a constant drain. Beyond the monthly line items and the continuous cost optimisation of running infrastructure, the primary driver for me is data containment. Ensuring that not a single packet of proprietary code or infrastructure configuration leaks into an outbound telemetry stream to an external corporate LLM provider is a massive win. As a follow-up to my last post exploring the trade-offs of RSS feeds versus heavy-handed cloud AI agents, I decided to run an architectural experiment. Can you build a complete, functional smartwatch application end to end inside a strictly self-contained, local AI stack? ...