Hardware & Semiconductors
184 articles RSS
Cerebras to Supply About 100 Megawatts of CS-4 Systems to Gimlet Labs for a Mixed-Silicon Inference Cloud
Cerebras will supply roughly 100 megawatts of CS-4 systems to inference-cloud startup Gimlet Labs, which pairs wafer-scale chips with GPUs and targets up to 3,000 tokens per second.
AWS Interconnect Adds Microsoft Azure in Public Preview, Extending Its Multicloud Networking Spec to a Third Hyperscaler
AWS and Microsoft launched a public preview letting customers privately link AWS and Azure workloads at up to 100 Gbps, completing the open interconnect spec's rollout across Azure, OCI, and Google Cloud.
Microsoft Details Maia 200 AI Accelerator Architecture at Hot Chips 2026
At Hot Chips 2026, Microsoft and an accompanying arXiv paper detailed Maia 200's software-defined dataflow architecture, a 750-watt inference chip delivering 10,145 Tflop/s of FP4 compute.
Anthropic Held, Then Abandoned, a $7 Billion Bid for AI Chip Startup MatX, Reuters Reports
Anthropic discussed buying AI chip startup MatX for roughly $7 billion, then walked away; the two sides are now said to be exploring a supply partnership instead.
SiFive Launches BigSky SF-2U870, a Rack-Mount RISC-V Server Now Running Nvidia CUDA
SiFive's BigSky SF-2U870 pairs 32 P870-D cores with CUDA support and Nvidia NVLink Fusion, targeting AI workload porting to RISC-V datacenters.
Apple's M6 and M5 Ultra Chips Debut a New Core AI Framework for On-Device Model Training
Apple's new M6 and M5 Ultra chips ship alongside Core AI, a new framework for building and deploying AI models on Apple silicon, with M5 Ultra supporting 512GB of unified memory.
NVIDIA Groq 3 LPX Inference Chip Enters Full Production, Claiming 4x Faster Response for Coding Agents
NVIDIA's Groq-derived Groq 3 LPX inference accelerator is now shipping, promising ultrafast token generation for agentic coding workloads, with Nebius first to deploy it.
Fractile Seeks $6.5 Billion Valuation After $250 Million Anthropic Chip Deal, Six Times Its May Price
UK inference-chip startup Fractile is in talks to raise about $600M at a $6.5B pre-money valuation, driven by an initial $250M chip deal with Anthropic.
Cerebras Launches CS-4 AI Accelerator, Claiming 30x Faster Inference Than GPUs on an Overclocked WSE-3
Cerebras unveiled its CS-4 rack-scale inference system, claiming 30x faster performance than GPUs, though independent analysis finds the chip inside is an overclocked WSE-3, not a new design.
NVIDIA's JetPack 7.2.1 Lets Developers Emulate the New Jetson T3000 Robotics Module on Existing Thor Hardware
JetPack 7.2.1 adds T3000 emulation on the Jetson Thor AGX Developer Kit plus new automated video-pipeline tools built on PyNvVideoCodec 2.2.
AMD's CDNA 5 Architecture Detailed: Inside the Instinct MI455X's 320-Billion-Transistor Design
A ServeTheHome deep dive breaks down CDNA 5, the architecture behind AMD's Instinct MI455X, alongside the ROCm.ai tooling meant to let AI coding agents program AMD GPUs natively.
AMD Acquires Taalas, Betting Model-Specific Chips That Etch Weights Into Silicon Can Boost AI Inference Speed
AMD is buying Toronto-based Taalas, whose chips etch model weights directly into silicon instead of storing them in memory, to strengthen its AI inference lineup.