Tagged28 Articles

#m5-max

Posts tagged with #m5-max from the Contra Collective team.

All Posts #AI104#Commerce73#Engineering69#Analytics5#Identity6#Strategy31#cache-invalidation1#headless-commerce28#isr1#edge-cache1#shopify-plus26#gemini-3-5-pro11#gpt-5-6-sol3#ifeval1#instruction-following1#model-comparison5#imatrix1#gguf1#quantization4#apple-silicon34#m5-max28#apollo-federation1#hasura1#wundergraph1#graphql1#whisper2#parakeet2#mlx29#speech-to-text1#opus-4-84#swe-bench-multimodal1#frontend2#product-feed1#feedonomics1#channable1#headless4#grok-4-51#tau-bench2#agentic-tools1#moe2#memory2#arc-agi-21#gpt-5.6-sol1#gemini-3.5-pro4#grok-4.54#reasoning4#benchmarks9#shopify-markets1#multi-region1#i18n1#internationalization1#unified-memory2#local-inference18#bigcodebench1#coding-benchmarks4#netsuite3#acumatica1#erp1#integration1#gpt-5.62#opus-4.84#gpqa-diamond3#perplexity1#subscriptions2#recharge2#ordergroove1#stripe-billing1#avalara2#taxjar2#vertex1#sales-tax2#tax-compliance2#fable-53#agents-last-exam1#pricing3#apple-neural-engine1#ane1#gpu1#core-ml1#swe-bench-pro4#terminal-bench4#agentic-coding16#sonnet-51#tokenizer1#throughput3#production1#gpt-5.54#prompt-caching2#cost1#claude-opus-4.81#qwen3-coder2#webdev-arena1#code-generation3#lora2#qlora1#fine-tuning2#bigcommerce1#b2b1#wholesale1#erp-integration1#claude-fable-57#grok-4.31#aime1#math-reasoning1#flux1#sdxl1#stable-diffusion1#image-generation1#nosto1#dynamic-yield1#rebuy1#personalization2#merchandising1#tool-use2#agents1#function-calling3#kokoro1#piper1#xtts1#text-to-speech1#rise-ai1#govalo1#gift-cards1#store-credit1#gpt-5-510#simpleqa1#hallucination1#factuality1#frontier-models12#rtx-50901#dgx-spark1#nvidia1#cost-per-token1#signifyd1#riskified1#nofraud1#fraud-prevention1#chargeback1#text-to-sql1#bird-benchmark1#spider-21#rag4#embeddings2#reranker2#yotpo1#okendo1#junip1#product-reviews1#ugc1#aider-polyglot3#livecodebench2#launchdarkly1#statsig1#growthbook1#feature-flags1#experimentation1#llama-cpp7#cold-start1#model-load1#first-token-latency1#claude-computer-use1#openai-operator1#perplexity-computer1#osworld1#webarena1#agentic-browsing1#loop-returns1#aftership1#narvar1#returns1#post-purchase1#m5-ultra6#m4-max1#tokens-per-watt1#energy-cost1#claude-sonnet-52#humanitys-last-exam1#research1#claude-opus-4-810#terminal-bench-21#fastly1#cloudflare-workers1#vercel-edge1#shopify-hydrogen3#edge-runtime1#vision1#docvqa1#chartqa1#document-understanding1#distil-whisper1#transcription1#asr1#long-context4#needle-in-haystack1#kv-cache3#nvme1#medusa2#vendure2#saleor2#migration1#structured-outputs2#tool-calling1#qwen32#llama-3-32#customer-support1#haiku-4-51#agent-cost1#triage1#qwen3-235b1#llama-3-3-70b2#mac-studio3#storyblok2#sanity3#headless-cms2#visual-editing1#content-ops1#webhooks1#stripe2#reliability1#idempotency1#claude-code2#codex-cli1#aider2#repo-migration1#cli1#fp81#production-inference1#salsify1#akeneo2#inriver1#pim2#algolia3#typesense3#meilisearch2#search3#bge1#cohere1#claude-haiku1#fable-5-mini1#gemini-3-1-flash2#cheap-tier2#cursor3#windsurf3#zed1#agentic-ide1#klaviyo2#iterable1#customer-io1#lifecycle-messaging1#order-management1#manhattan-active-omni1#fluent-commerce1#oms1#speculative-decoding1#draft-model1#claude-sonnet-4-63#swe-bench4#searchspring1#constructor-io1#klevu1#ecommerce-search1#skio1#smartrr1#bold1#local-agents1#outlines1#claude-haiku-4-52#gemini-3-5-flash1#swe-bench-verified1#bloomreach1#attentive1#cdp2#reasoning-benchmarks1#prefix-cache1#contentful1#composable-commerce1#shipstation1#shipengine1#easypost1#fulfillment1#shipping-api1#batched-inference1#multi-user1#anthropic2#stripe-tax1#grok-4-32#m5-pro3#decode-throughput1#shopify-payments1#payments1#checkout1#enterprise-ecommerce1#pimcore1#plytix1#product-information-management1#salesforce-commerce-cloud1#m4-pro2#power-consumption1#tco1#joules-per-token1#okta1#auth01#enterprise-identity1#iam1#security2#gpt-5-5-mini1#llm-comparison4#model-pricing1#mlx-lm1#prefill1#decode1#batching1#posthog1#amplitude2#product-analytics2#event-tracking1#sustained-throughput1#thermal-throttling1#ide2#ai-coding2#cloudinary1#imgix1#shopify-cdn1#image-cdn1#performance1#hydrogen1#nextjs1#vision-llm1#qwen-vl1#moondream1#llava1#multimodal1#gemini-3-1-pro1#open-source1#sfcc1#aws-lambda1#cloud-run1#serverless2#ai-inference1#infrastructure23#ai-models2#claude-mythos-51#ai-policy1#export-controls1#ai-governance1#flash-attention1#local-llm2#ai-infrastructure6#redis1#upstash1#caching1#ai-search1#distributed-inference1#claude1#fable1#opus1#gpt1#gemini1#llm1#github-actions1#gitlab-ci1#ci-cd2#devops2#mixpanel1#data1#ecommerce3#segment1#rudderstack1#crewai1#autogen1#multi-agent1#ai-agents1#enterprise-ai2#datadog1#new-relic1#observability1#apm1#langfuse2#helicone1#llm-observability2#ai-monitoring1#langsmith1#evaluation1#langchain1#supabase1#firebase1#baas1#backend2#ai-apps1#llamaindex1#haystack1#nlp1#wandb1#mlflow1#mlops1#experiment-tracking1#aws-bedrock1#vertex-ai1#managed-llm1#cloud1#developer-tools1#models1#tools1#nodejs1#supply-chain1#npm1#pypi1#tanstack1
Jun 17, 2026AI Infrastructure

Joules per Token on Apple Silicon: Local LLM Power and TCO (M4 Pro / M5 Pro / M5 Max / M5 Ultra, 2026)

Tokens per second is half the story. The other half is what those tokens cost in electricity, hardware amortization, and rack space. Most Apple Silicon local LLM benchmarks publish a decode number on Llama 3.1 8B and stop. That makes for a clean headline and a useless TCO model. We measured wall power on M4 Pro, M5 Pro, M5 Max, and M5 Ultra under sustained inference for Llama 3.3 8B and Llama 3.3 70B, then converted everything to Joules per generated token and dollars per million tokens. The numbers settle the "should we self host on Apple Silicon" conversation in a way that throughput alone cannot.

Jun 16, 2026AI Infrastructure

mlx-lm vs llama.cpp on M5 Max: Prefill, Decode, and Batch Scaling for 70B (2026)

Most Apple Silicon local inference benchmarks publish a single decode number for an 8B parameter model and stop. That works for a demo and falls apart the moment you try to put MLX or llama.cpp behind a real backend. The interesting question is what prefill latency, decode throughput, and batch size scaling actually look like on a 70B class model under the two runtimes most teams pick between in mid 2026: mlx-lm and llama.cpp with the Metal backend.

Jun 15, 2026AI Infrastructure

Apple Silicon Sustained LLM Throughput: M4 Pro vs M5 Pro vs M5 Max Thermal Tested (2026)

Every Apple Silicon local inference benchmark on the internet runs for thirty seconds, quotes a peak token rate, and calls it a day. That number is useful for marketing copy and useless for capacity planning. If you are deploying MLX or llama.cpp behind a real backend, the question is not what the SoC does in the first thirty seconds. The question is what it does over four hours of continuous decode under realistic batch conditions, after the package power has saturated and the thermal envelope has clamped down.