Intel Core Ultra
Embedded NPU acceleration for lightweight, low-latency inference such as inspection, maintenance, and local monitoring.
Place AI, LLM inference, and other operational workloads across eligible fabric locations to support lower-latency decisions without separating edge compute from network governance.
Operational value
Selected Service Endlets can run approved inference close to operational data. Placement remains governed by locality, latency, sovereignty, capacity, policy, and execution requirements, with the surrounding fabric providing protected connectivity and operating context.
Inside the capability
Selected Service Endlets can run eligible AI and LLM inference closer to operations, devices, and data. Placement policy accounts for locality, latency, sovereignty, capacity, and the approved execution environment.
Embedded NPU acceleration for lightweight, low-latency inference such as inspection, maintenance, and local monitoring.
AMX acceleration for approved enterprise inference workloads, RAG systems, and autonomous agents.
Inference remains inside the Endlet identity, policy, connectivity, telemetry, and lifecycle model.
The Endlet operating model
Why it matters
Place eligible inference closer to machines, sensors, users, and operational data.
Support architectures that process selected information near its source when policy requires it.
Coordinate workload placement with identity, connectivity, telemetry, and lifecycle controls.
How it operates
Establish locality, latency, sovereignty, capacity, and execution constraints.
Place the approved workload on a suitable Service Endlet close to operations.
Maintain fabric telemetry, identity, policy, and lifecycle context around inference.
Coordinated by design
Design your fabric
Bring your current topology, providers, workloads and continuity requirements. We’ll map the 21Packets capabilities that fit.
Book a working session