Tagged6 Articles

#m5-ultra

Posts tagged with #m5-ultra from the Contra Collective team.

All Posts #AI104#Commerce73#Engineering69#Analytics5#Identity6#Strategy31#cache-invalidation1#headless-commerce28#isr1#edge-cache1#shopify-plus26#gemini-3-5-pro11#gpt-5-6-sol3#ifeval1#instruction-following1#model-comparison5#imatrix1#gguf1#quantization4#apple-silicon34#m5-max28#apollo-federation1#hasura1#wundergraph1#graphql1#whisper2#parakeet2#mlx29#speech-to-text1#opus-4-84#swe-bench-multimodal1#frontend2#product-feed1#feedonomics1#channable1#headless4#grok-4-51#tau-bench2#agentic-tools1#moe2#memory2#arc-agi-21#gpt-5.6-sol1#gemini-3.5-pro4#grok-4.54#reasoning4#benchmarks9#shopify-markets1#multi-region1#i18n1#internationalization1#unified-memory2#local-inference18#bigcodebench1#coding-benchmarks4#netsuite3#acumatica1#erp1#integration1#gpt-5.62#opus-4.84#gpqa-diamond3#perplexity1#subscriptions2#recharge2#ordergroove1#stripe-billing1#avalara2#taxjar2#vertex1#sales-tax2#tax-compliance2#fable-53#agents-last-exam1#pricing3#apple-neural-engine1#ane1#gpu1#core-ml1#swe-bench-pro4#terminal-bench4#agentic-coding16#sonnet-51#tokenizer1#throughput3#production1#gpt-5.54#prompt-caching2#cost1#claude-opus-4.81#qwen3-coder2#webdev-arena1#code-generation3#lora2#qlora1#fine-tuning2#bigcommerce1#b2b1#wholesale1#erp-integration1#claude-fable-57#grok-4.31#aime1#math-reasoning1#flux1#sdxl1#stable-diffusion1#image-generation1#nosto1#dynamic-yield1#rebuy1#personalization2#merchandising1#tool-use2#agents1#function-calling3#kokoro1#piper1#xtts1#text-to-speech1#rise-ai1#govalo1#gift-cards1#store-credit1#gpt-5-510#simpleqa1#hallucination1#factuality1#frontier-models12#rtx-50901#dgx-spark1#nvidia1#cost-per-token1#signifyd1#riskified1#nofraud1#fraud-prevention1#chargeback1#text-to-sql1#bird-benchmark1#spider-21#rag4#embeddings2#reranker2#yotpo1#okendo1#junip1#product-reviews1#ugc1#aider-polyglot3#livecodebench2#launchdarkly1#statsig1#growthbook1#feature-flags1#experimentation1#llama-cpp7#cold-start1#model-load1#first-token-latency1#claude-computer-use1#openai-operator1#perplexity-computer1#osworld1#webarena1#agentic-browsing1#loop-returns1#aftership1#narvar1#returns1#post-purchase1#m5-ultra6#m4-max1#tokens-per-watt1#energy-cost1#claude-sonnet-52#humanitys-last-exam1#research1#claude-opus-4-810#terminal-bench-21#fastly1#cloudflare-workers1#vercel-edge1#shopify-hydrogen3#edge-runtime1#vision1#docvqa1#chartqa1#document-understanding1#distil-whisper1#transcription1#asr1#long-context4#needle-in-haystack1#kv-cache3#nvme1#medusa2#vendure2#saleor2#migration1#structured-outputs2#tool-calling1#qwen32#llama-3-32#customer-support1#haiku-4-51#agent-cost1#triage1#qwen3-235b1#llama-3-3-70b2#mac-studio3#storyblok2#sanity3#headless-cms2#visual-editing1#content-ops1#webhooks1#stripe2#reliability1#idempotency1#claude-code2#codex-cli1#aider2#repo-migration1#cli1#fp81#production-inference1#salsify1#akeneo2#inriver1#pim2#algolia3#typesense3#meilisearch2#search3#bge1#cohere1#claude-haiku1#fable-5-mini1#gemini-3-1-flash2#cheap-tier2#cursor3#windsurf3#zed1#agentic-ide1#klaviyo2#iterable1#customer-io1#lifecycle-messaging1#order-management1#manhattan-active-omni1#fluent-commerce1#oms1#speculative-decoding1#draft-model1#claude-sonnet-4-63#swe-bench4#searchspring1#constructor-io1#klevu1#ecommerce-search1#skio1#smartrr1#bold1#local-agents1#outlines1#claude-haiku-4-52#gemini-3-5-flash1#swe-bench-verified1#bloomreach1#attentive1#cdp2#reasoning-benchmarks1#prefix-cache1#contentful1#composable-commerce1#shipstation1#shipengine1#easypost1#fulfillment1#shipping-api1#batched-inference1#multi-user1#anthropic2#stripe-tax1#grok-4-32#m5-pro3#decode-throughput1#shopify-payments1#payments1#checkout1#enterprise-ecommerce1#pimcore1#plytix1#product-information-management1#salesforce-commerce-cloud1#m4-pro2#power-consumption1#tco1#joules-per-token1#okta1#auth01#enterprise-identity1#iam1#security2#gpt-5-5-mini1#llm-comparison4#model-pricing1#mlx-lm1#prefill1#decode1#batching1#posthog1#amplitude2#product-analytics2#event-tracking1#sustained-throughput1#thermal-throttling1#ide2#ai-coding2#cloudinary1#imgix1#shopify-cdn1#image-cdn1#performance1#hydrogen1#nextjs1#vision-llm1#qwen-vl1#moondream1#llava1#multimodal1#gemini-3-1-pro1#open-source1#sfcc1#aws-lambda1#cloud-run1#serverless2#ai-inference1#infrastructure23#ai-models2#claude-mythos-51#ai-policy1#export-controls1#ai-governance1#flash-attention1#local-llm2#ai-infrastructure6#redis1#upstash1#caching1#ai-search1#distributed-inference1#claude1#fable1#opus1#gpt1#gemini1#llm1#github-actions1#gitlab-ci1#ci-cd2#devops2#mixpanel1#data1#ecommerce3#segment1#rudderstack1#crewai1#autogen1#multi-agent1#ai-agents1#enterprise-ai2#datadog1#new-relic1#observability1#apm1#langfuse2#helicone1#llm-observability2#ai-monitoring1#langsmith1#evaluation1#langchain1#supabase1#firebase1#baas1#backend2#ai-apps1#llamaindex1#haystack1#nlp1#wandb1#mlflow1#mlops1#experiment-tracking1#aws-bedrock1#vertex-ai1#managed-llm1#cloud1#developer-tools1#models1#tools1#nodejs1#supply-chain1#npm1#pypi1#tanstack1
Jun 17, 2026AI Infrastructure

Joules per Token on Apple Silicon: Local LLM Power and TCO (M4 Pro / M5 Pro / M5 Max / M5 Ultra, 2026)

Tokens per second is half the story. The other half is what those tokens cost in electricity, hardware amortization, and rack space. Most Apple Silicon local LLM benchmarks publish a decode number on Llama 3.1 8B and stop. That makes for a clean headline and a useless TCO model. We measured wall power on M4 Pro, M5 Pro, M5 Max, and M5 Ultra under sustained inference for Llama 3.3 8B and Llama 3.3 70B, then converted everything to Joules per generated token and dollars per million tokens. The numbers settle the "should we self host on Apple Silicon" conversation in a way that throughput alone cannot.