Frontier Technology Portal Independent technology analysis / Updated daily
Frontier Technology Portal logo
FRONTIER Technology Portal for the next wave of invention

Category: Artificial Intelligence

Clear explainers about AI models, agents, chips, automation, and responsible deployment.

  • How Speculative Decoding Speeds Up AI Text Generation

    How Speculative Decoding Speeds Up AI Text Generation

    A chatbot’s answer often arrives a few words at a time. Behind that familiar stream is a costly sequence of model calculations. Speculative decoding tries to shorten that sequence: a fast assistant proposes several tokens, and the main model checks them together. When enough proposals survive, the system advances farther with each expensive verification step.

    The appeal is faster output without simply replacing the model you asked for with a smaller one. But the details matter. Verification here means checking a proposed continuation against a model’s probability distribution, not checking whether the answer is true. And a technique that speeds up one conversation can be less helpful on a busy server.

    Why generating an answer can be slow

    An autoregressive language model predicts the next token using the text already available. A token might be a word, part of a word, or punctuation. After producing one, ordinary decoding repeats the process with that token added to the context. This dependency makes the output stage difficult to parallelize in the straightforward way used for processing many already-known tokens.

    The original ICML paper on speculative decoding tackled this serial bottleneck by letting a cheaper model do provisional work. The larger model still controls the accepted output. This is a change to the generation procedure, not a promise that a small model has suddenly acquired the larger model’s capabilities.

    It also does not eliminate the cost of reading a long prompt or storing context. Our explanation of long context windows and KV-cache memory covers a separate pressure that remains relevant even when output tokens arrive faster.

    Draft several tokens, then verify the sequence

    Imagine the assistant proposing a short continuation. The target model can score those proposed positions in one forward pass. The acceptance procedure works from left to right; it does not freely assemble an answer from whichever later tokens look appealing.

    For sampled generation, exact verification is more sophisticated than asking whether the models chose the same favorite word. The draft and target probabilities enter a rejection-sampling procedure. If a proposal is rejected, the procedure supplies a correction and continues from the accepted prefix rather than trusting the rejected continuation.

    The original speculative sampling research explains how this can recover the target model’s distribution while allowing multiple tokens to pass through one target evaluation. The savings depend on the draft being both cheap and useful. A fast assistant that repeatedly suggests unsuitable continuations creates extra work.

    Preserving a distribution is not preserving every answer

    In the exact formulation, speculative decoding is designed to preserve what the target model would sample statistically. That is different from guaranteeing that every run produces the same sentence. Random sampling already permits different valid outputs, and numerical behavior adds another complication.

    The vLLM documentation describes theoretical losslessness up to hardware numerical precision. It also warns that floating-point behavior and batch changes can affect probabilities and outputs. A blanket claim of identical text under every deployment setting goes beyond that guarantee.

    Most importantly, a target model can confidently accept an inaccurate answer. This mechanism does not inspect a scientific paper, authenticate a citation, or consult a database. Retrieval-augmented generation and grounding evaluation address a different question: what evidence an answer uses, and whether that evidence actually supports it.

    Why a longer draft is not always better

    Successful proposals let the system skip some target-generation steps, but drafting and verifying them still consume resources. If a long block is rejected near its beginning, much of that work does not advance the answer. Draft length therefore needs to be considered alongside acceptance rate, not advertised as a speed control by itself.

    Load changes the tradeoff too. With a small batch, a server can be constrained by moving model data rather than performing arithmetic. At larger batches, the same hardware may become more compute-bound, making extra speculative calculations less attractive.

    vLLM’s current adaptive verification documentation explicitly discusses this distinction. Its adaptive implementation chooses draft lengths using confidence estimates and a profiled compute budget. The documented feature currently requires DSpark models with a confidence head and is disabled by default; it is not a universal switch for every model.

    An assistant model is not the only drafting option

    The familiar setup uses a separate small model, but the broader idea has other implementations. Hugging Face documents prompt-lookup decoding, which proposes repeated sequences from the input itself. That can be useful when an output reuses source text, such as in some summarization or translation tasks, without loading another drafting model.

    It also documents self-speculative decoding using early layers of a model trained to support early-exit predictions. Tokenizer compatibility needs attention: conventional assistant arrangements commonly share a tokenizer, while universal assisted decoding converts between representations and handles alignment.

    These options in the Transformers assisted-decoding guide are implementation choices, not interchangeable guarantees. For someone trying a local model, the useful first question is whether the selected model and serving software support the intended method, rather than whether the model is described broadly as “speculative-ready.”

    What a meaningful speed comparison should show

    A smooth-looking demo is not enough. Separate the wait before the first token from the pace of subsequent output. Then distinguish the experience of one user from total throughput across many users. The MLCommons endpoint benchmark makes this distinction between per-user interactivity and aggregate throughput explicit.

    A useful comparison should report the target and drafting setup, hardware, concurrency, prompt lengths, output lengths, sampling settings, and the non-speculative baseline. Ask whether the reported result is a quiet single-user session or a loaded service. Acceptance rate helps explain a result, but cannot replace end-to-end timing or resource measurements.

    MLPerf Inference v6.0 introduced a standardized speculative-decoding configuration for its DeepSeek-R1 Interactive workload. The MLCommons announcement specifies the official multi-token-prediction head with EAGLE-style decoding. That controlled benchmark is more interpretable than an unspecified speed headline, but it does not validate every assistant-model pairing.

    Limitations and what to watch next

    Speculative decoding is best understood as an opportunity to use otherwise expensive generation steps more efficiently. It is not a substitute for enough memory, an evidence-based answer, or a properly measured deployment. This article explains published methods and documentation; it does not report hands-on performance testing.

    For a local enthusiast, a sensible trial compares the same target model on the same representative prompts with speculation enabled and disabled. Measure responsiveness and memory use, and retain a fallback configuration. Model quantization changes numerical representation; speculation changes how proposed output is checked. Results from combining them should be measured as a combination, not attributed automatically to either technique.

    Watch for better workload-aware controls, clearer compatibility documentation, and benchmarks that report both latency and throughput. The important advance will not be the longest possible draft. It will be a system that knows when a draft is likely to save useful time, and when ordinary decoding is the better choice.

    Primary and authoritative sources

    Featured image is an AI-generated editorial metaphor, not a photograph of actual speculative-decoding hardware.

  • Long AI Context Windows Have a Memory Bill

    Long AI Context Windows Have a Memory Bill

    A large language model can advertise a very long context window and still struggle to serve many long conversations at once. The model’s weights may fit comfortably on an accelerator, yet the server can run short of memory as prompts and generated answers grow. The missing part of the explanation is the key-value cache, usually shortened to KV cache.

    The KV cache is temporary inference state. It keeps intermediate attention data for tokens the model has already processed, allowing each new token to reuse earlier work. That reuse makes autoregressive generation practical, but the saved state consumes memory for every layer, every active sequence, and every token retained in context. A longer window is therefore not only a model capability; it is a serving cost.

    Generation has a prefill phase and a decode phase

    When a prompt first arrives, the model processes many input tokens together. This is the prefill phase. Modern accelerators can exploit substantial parallelism here, although a very long prompt still requires considerable computation and memory traffic.

    After prefill, the model enters decode. It produces one token, adds that token to the context, then produces the next. Each step depends on the earlier sequence. Recalculating all earlier attention projections at every step would repeat a large amount of work.

    The current Hugging Face Transformers cache documentation explains the solution: store the key and value vectors created for previous tokens and reuse them during later steps. Only the new token’s key and value need to be added. This reduces repeated computation, but memory usage grows as the sequence grows.

    What the cache actually stores

    Transformer attention creates query, key, and value vectors. A new query is compared with keys from earlier tokens to determine which stored values should influence the next representation. During causal generation, the keys and values for completed tokens do not need to be recomputed, so they can be cached.

    The rough memory relationship is straightforward: cache size grows with the number of model layers, retained tokens, key-value heads, head dimension, bytes per stored element, and simultaneous sequences. There are two tensors, keys and values, so both count. The model’s parameter count can remain unchanged while runtime memory rises linearly with context length and batch size.

    This is different from weight memory. Techniques discussed in our guide to AI model quantization can shrink model weights, yet a service may still face a large KV cache. Some systems also quantize the cache, but that is a separate choice with its own performance and quality considerations.

    A context-window limit is not a concurrency guarantee

    A provider may be able to run one request near the model’s maximum window but not hundreds of such requests on the same hardware. Every active conversation reserves or consumes cache capacity. Longer prompts can reduce the number of requests that fit together, which can lower throughput or increase queueing.

    Generation length matters too. A prompt can begin well below the advertised limit and approach it as the answer grows. Beam search or requests for multiple candidate answers can multiply active sequence state. Multimodal inputs may also become many tokens before attention begins.

    This is why an impressive maximum window does not describe typical service speed. Useful measurements include time to first token, tokens generated per second, throughput under concurrent load, and behavior at several prompt lengths. The same principle applies to on-device AI benchmarks: a headline capacity is not a complete workload test.

    Multi-query and grouped-query attention reduce the bill

    Traditional multi-head attention can maintain separate key and value projections for many attention heads. The original multi-query attention paper proposed sharing one set of keys and values across query heads. That sharply reduces the amount of KV data that must be stored and moved during incremental decoding.

    Sharing everything can trade away some model quality or flexibility. Grouped-query attention, or GQA, offers a middle ground: groups of query heads share a smaller number of key-value heads. The original GQA study describes it as an interpolation between full multi-head attention and multi-query attention, aiming for quality closer to the former and speed closer to the latter.

    These are architectural choices, not switches that can always be added after deployment. They change how the model is trained or adapted and how its cache is shaped.

    PagedAttention tackles wasted space

    Even when the cache data itself is necessary, memory can be wasted through allocation. Requests arrive with different prompt lengths and stop at unpredictable times. Reserving one large contiguous region for each sequence can leave unused capacity and fragmentation.

    The PagedAttention paper applied an operating-system idea to this problem. It divides KV cache data into fixed-size blocks that do not need to occupy one contiguous region. A serving engine can allocate blocks as a sequence grows, share blocks where appropriate, and reclaim them when requests finish.

    Paging does not eliminate the cache or make long context free. It improves memory utilization and scheduling, allowing hardware to host more useful work before capacity is exhausted. NVIDIA’s current TensorRT-LLM documentation lists paged KV caching alongside in-flight batching and quantization as production inference optimizations.

    Prefix reuse can avoid repeated prefill work

    Many requests begin with an identical system prompt, policy document, or application template. Prefix caching stores the KV state for that shared beginning so later requests can reuse it instead of running the same prefill again. This can reduce latency and compute when the prefix is exactly reusable.

    The match must be precise. A changed token, different model revision, altered positional treatment, or incompatible cache format can prevent reuse. Cached prefixes also occupy memory and require eviction rules. The optimization is most valuable when a stable prefix is long and requested frequently.

    Offloading, quantization, and sliding windows trade something

    When accelerator memory is tight, a runtime can move some cache data to host memory. That frees scarce device memory but adds transfers over a slower link. Cache quantization stores each element with fewer bits, reducing capacity and bandwidth demands while introducing additional conversion work and possible numerical effects.

    Sliding-window attention limits how far selected layers look back, so their cache stops growing after a defined range. Other systems summarize, compress, or discard older state. These methods control memory, but they are not identical to full attention over every retained token. Product claims should distinguish an accepted input length from the effective attention pattern used across that input.

    Retrieval can reduce the need to paste an entire library into every prompt by selecting relevant passages first. But as our RAG guide explains, retrieval adds ranking and source-quality problems; it is not a free substitute for reasoning over context.

    Long context also affects energy and cooling

    More active memory, longer attention operations, and lower batching efficiency can increase hardware time per request. The exact energy impact depends on model architecture, hardware, software, utilization, and workload. It cannot be inferred from token count alone.

    At data-center scale, serving design connects directly to power delivery and heat removal. Our article on AI server cooling architecture explains why thermal systems must be designed around sustained component power and rack density, not only a chip’s peak rating.

    Limitations and what to watch next

    KV cache behavior varies by architecture. Full attention, sliding-window attention, hybrid models, recurrent state-space layers, multimodal encoders, and speculative decoding can create different memory patterns. Vendor APIs rarely expose enough detail to predict cost from a context-window number alone.

    Watch for clearer reporting of latency and throughput across prompt lengths, wider use of cache-aware schedulers, better low-bit cache formats, and architectures that reduce state without losing useful long-range information. For users, the practical lesson is simple: request only the context that improves the answer. For operators, long context must be budgeted as live memory, bandwidth, and concurrency, not treated as a line on a model card.

    Featured image: AI-generated editorial visualization of a growing KV cache inside an unbranded inference server. It is not a photograph of a specific product or a hands-on hardware test.

    Primary and authoritative sources

  • AI Agents Need Their Own Identities, Not Your Login

    AI Agents Need Their Own Identities, Not Your Login

    An AI agent that can read email, update a calendar, edit files, or call business software needs permission to act. The fastest setup is often the most dangerous: give the agent a person’s login, paste in a long-lived API key, and let it operate as if it were the user. That may make a demonstration work, but it also erases an important boundary. A service can no longer reliably tell whether the person acted, the agent acted, or someone stole the shared credential.

    In August 2026, the US National Institute of Standards and Technology argued that agents should be treated as first-class digital entities with their own identifiers, credentials, and entitlements. This is less glamorous than a new reasoning benchmark, but it may matter more to everyday adoption. Agents that take actions need an identity and authorization layer designed for software moving at machine speed.

    Borrowing a user’s login destroys accountability

    Authentication answers who or what is making a request. Authorization answers what that identity may do. Delegation connects the two by recording that a person or organization allowed an agent to perform a defined task on its behalf.

    If an agent simply uses a human account, all three ideas collapse into one shared login. Audit logs attribute its actions to the person. Revoking the agent may require changing the person’s credential. Permission settings cannot easily distinguish between actions the user may take directly and the smaller set the agent actually needs.

    NIST’s August 2026 identity guidance warns that credential sharing creates security, privacy, legal, and non-repudiation problems. Its preferred direction is a unique agent identity whose rights are bound to the user or system that operates it. The audit trail can then say that a particular agent acted under a particular delegation, rather than pretending the human clicked every button.

    An agent identity is not a second human account

    Creating a separate username for an agent is a start, not a complete architecture. The identity should describe a software workload and its operating context: who owns it, which service runs it, which version or deployment is active, what task it was assigned, and when its authority expires. A short-lived agent should not inherit a credential that remains valid for months.

    The same distinction matters when agents call other agents. Our guide to AI agent protocols explains how systems exchange tools and messages. Interoperability does not establish permission. Every hop needs a verifiable chain showing which identity delegated which authority, resource, and purpose.

    NIST launched an AI Agent Standards Initiative in 2026 around secure interoperability, identity research, open protocols, and evaluation. The related NCCoE project is reviewing how existing identity standards can be applied to enterprise agents. The effort is still developing, so product claims about a universal agent identity standard deserve caution.

    Long-lived bearer tokens are a poor default

    Many integrations use bearer tokens: whoever possesses the token can present it to an API. That simplicity becomes risky when an agent moves across prompts, tools, plug-ins, logs, and network services. A token accidentally written to a configuration file or transcript may be reusable by anyone who finds it.

    The problem is magnified by lifespan and scope. A key that can read and delete every file is far more dangerous than a token that can read one folder for ten minutes. An agent may also explore several routes to complete a broad instruction, exposing credentials to more components than its designer expected.

    The IETF’s OAuth 2.0 security best practice recommends restricting an access token to the minimum privileges required, limiting it to intended resource servers, and binding it to a particular sender where possible. These are general web security principles, not agent-specific inventions. Agentic systems make them more urgent because one vague instruction can trigger many rapid requests.

    Useful authorization describes the task, not just the role

    A coarse role such as “editor” or “administrator” is often wider than an automated task requires. Better authorization can name the permitted resource and action: create a draft but do not publish it; read invoices from one account but do not change payment details; schedule a meeting inside working hours but do not invite external guests.

    Time, spending limits, data sensitivity, destination, and required review can all be policy inputs. Rights should narrow as a task is delegated through a chain, rather than expanding at each step. When the task ends, its temporary authority should end too. This approach reduces damage from model errors, prompt injection, a compromised tool, or an agent that chooses an unexpected path.

    Authorization cannot replace application safety. File services still need recovery, payment systems need fraud controls, and production infrastructure needs change review. Identity establishes who is asking; policy and resilient design determine acceptable consequences.

    Proof of possession can limit stolen-token reuse

    A sender-constrained token is usable only when the client can also prove control of associated key material. Stealing the token alone is therefore less useful. The IETF’s DPoP standard defines an application-level proof-of-possession mechanism for OAuth tokens. Each relevant request carries a signed proof tied to the client’s key.

    This is not a magic shield. If an attacker controls the agent’s execution environment and can use both the token and private key, the protection may fail. Key storage, software isolation, token lifetime, logging, and resource-server checks still matter. Confidential computing can protect some data while it is being processed, but as our confidential AI explainer notes, no single mechanism secures the entire pipeline.

    Human approval can become another weak point

    Requiring approval before a consequential action is sensible, but asking too often creates consent fatigue. People learn to click allow just to keep work moving, much as repeated authentication prompts can condition users to approve a malicious request. NIST specifically highlights this risk for agents that repeatedly ask for new tools or data.

    An approval should therefore present a meaningful transaction: what will happen, which data and destination are involved, how much authority is being granted, and whether the action can be reversed. Routine low-risk work can operate within a preapproved policy. Irreversible, unusual, or high-impact actions can trigger a focused checkpoint. The goal is not maximum prompting; it is useful human control.

    What buyers and administrators should look for

    A serious agent platform should make separate identities visible in administration and audit tools. It should support short-lived credentials, narrow scopes, explicit resource audiences, rapid revocation, secret-free logs, and a record of the human or service that delegated authority. Administrators should be able to disable one agent without disabling its owner.

    Ask whether permissions are checked only when a task begins or again when the agent switches tools and destinations. Check how sub-agents inherit rights, where keys are stored, what happens after a failed action, and whether logs distinguish proposed, approved, attempted, and completed operations. A polished chat interface says little about these controls.

    NIST’s 2026 analysis of public comments on agent security found broad agreement that established cybersecurity practices remain relevant but need adaptation for agent systems. That is a useful frame: most building blocks already exist, yet their combination and operational scale are new.

    Limitations and what to watch next

    Agent identity will not make model outputs reliable, stop every prompt injection, or prove that an action matches a user’s true intent. Standards are still evolving, and consumer services face an especially difficult problem when an outside agent arrives with credentials copied from a user. Cross-company delegation, revocation, liability, and privacy remain unsettled.

    The next meaningful signals will be interoperable implementations, not new terminology. Watch for NIST’s practical guidance, tests that compare agent security behavior, standard ways to pass reduced authority through multi-agent chains, and service providers that accept agent-specific credentials without encouraging password sharing. Products should also demonstrate that revocation propagates quickly and that audit records survive across tool boundaries.

    An AI agent should be able to act for you without becoming indistinguishable from you. Giving it a separate identity, limited authority, and a traceable delegation chain is the foundation for that difference.

    Featured image: AI-generated editorial illustration of agent identity and authorization controls, not a screenshot of a specific commercial system.

    Primary sources

  • AI Quantization Shrinks Models, but It Does Not Guarantee Faster Inference

    AI Quantization Shrinks Models, but It Does Not Guarantee Faster Inference

    Modern AI models move enormous arrays of numbers between memory and compute units. Reducing the precision of those numbers can shrink a model, lower memory traffic, and unlock faster arithmetic. This process, called quantization, is one of the main reasons a model that once required a large server can sometimes run on a smaller accelerator or local device.

    Quantization is not a universal compression switch. The result depends on which tensors are converted, how their ranges are measured, what hardware instructions and kernels exist, and whether the lower-precision model still performs the intended task. A file can become smaller without inference becoming faster, and an average benchmark can hide rare but important quality failures.

    Quantization maps a large numeric range into fewer values

    Training commonly uses 32-bit or 16-bit floating-point numbers. Quantization represents some weights or activations with formats such as 8-bit integers, 4-bit integers, or shorter floating-point types. A scale, and sometimes a zero point, connects the compact values to the approximate original range.

    Fewer bits reduce storage. Moving fewer bytes can ease a memory-bandwidth bottleneck, and specialized hardware can perform more low-precision operations per cycle. The approximation also discards information. Whether that loss matters depends on the distribution of values, the model architecture, and the task.

    The current PyTorch torchao quantization overview separates algorithms, quantized tensor layouts, primitive operations, and efficient kernels. That layered view is useful: a quantized data type alone does not guarantee an optimized execution path.

    Weight-only quantization solves part of the problem

    Model weights are stored parameters reused for every request. Converting only the weights can sharply reduce the memory needed to load a large model while leaving activations in a higher-precision format. This is attractive when model size and memory bandwidth dominate.

    During inference, however, the runtime may need to unpack or dequantize weights before multiplication. If the accelerator lacks a fused kernel for the exact format and block size, conversion overhead can consume the expected gain. Small batches or layers not dominated by matrix multiplication may see little benefit.

    Quantizing activations as well can reduce more traffic and enable lower-precision matrix operations. Activations vary with each input and can contain difficult outliers, so this step is often more sensitive than compressing weights alone.

    Per-tensor, per-channel, and per-block scales trade metadata for fidelity

    One scale for an entire tensor is simple but forces small and large values into the same grid. Separate scales for each output channel or small block can match local ranges more closely. That usually preserves more information at the cost of extra scale data, more complicated packing, and stricter kernel requirements.

    Symmetric schemes use a range centered around zero; asymmetric schemes add a zero point to fit an uneven range. Granularity, group size, rounding, clipping, and numeric format are part of the model artifact. Calling two files INT4 does not mean they use the same representation or will run on the same engine.

    NVIDIA’s current TensorRT quantization documentation, updated in August 2026, lists different rules for INT8, FP8, block-scaled formats, INT4, and FP4. For example, its INT4 support is weight-only with specified block sizes. This illustrates how hardware software contracts shape practical deployment.

    Post-training quantization needs representative calibration

    Post-training quantization converts an already trained model. Dynamic quantization measures some activation ranges during inference, adapting to each input but adding work. Static quantization measures ranges in advance using a calibration dataset, then stores the chosen parameters.

    The official ONNX Runtime quantization guide describes both workflows and several static calibration methods. It also warns that performance depends on hardware support and that quantize-dequantize overhead can make a model slower on devices without suitable instructions.

    Calibration data should represent the deployment distribution, including long prompts, unusual images, accents, languages, sensor conditions, and other difficult inputs. A narrow sample may choose ranges that look excellent in a demo and clip values encountered in production.

    Outliers make transformer activations difficult

    A small number of activation channels in a transformer can be much larger than the rest. If one scale must cover those outliers, most ordinary values occupy only a small part of the quantized range. Clipping the outliers can improve resolution for typical values but may remove meaningful signal.

    The original SmoothQuant research proposed moving part of this quantization difficulty from activations into weights through a mathematically equivalent rescaling before conversion. Its experiments demonstrated that an algorithm can preserve quality while enabling efficient 8-bit weight-and-activation execution on supported systems.

    The broader lesson is not that one method solves every model. Outlier structure changes across architectures and layers, so quantization needs model-specific evaluation rather than a universal bit-width claim.

    Quantization-aware training can recover quality

    If post-training conversion loses too much accuracy, quantization-aware training simulates lower-precision rounding during training or fine-tuning. The model can adapt its weights to the constraints before final conversion.

    This can improve quality at an added cost: training data, compute, engineering, and a new validation cycle. It also creates a distinct model artifact that must be governed, tested, and documented. A quantized model is not merely the original checkpoint stored in a smaller archive.

    Unsupported operators can erase the speedup

    Real models include normalization, attention, embeddings, activation functions, routing, sampling, and data movement around large matrix multiplications. A runtime may execute supported layers at low precision but fall back to a higher precision for others. Conversions between formats add latency and memory traffic.

    Batch size, sequence length, prompt length, output length, and cache behavior also change the bottleneck. Weight compression may help initial model loading and token generation differently. The key-value cache used by transformer attention can become a major memory consumer during long-context or high-concurrency serving even when weights are compact.

    This is why our guide to AI chips, memory, and power treats data movement as part of compute rather than an afterthought.

    Quality must be measured on the intended behavior

    Perplexity or average task accuracy is a useful screen, but it may not reveal instruction-following failures, formatting errors, tool-call changes, multilingual degradation, calibration shifts, or rare safety-relevant mistakes. Generative outputs also vary, so matched prompts and statistical testing are important.

    Teams should compare the original and quantized model across representative tasks, long contexts, adversarial cases, and subgroups. They should record the quantization method, calibration data, runtime, kernel version, sampling settings, and hardware. The same compact file can behave differently after a runtime update.

    The need for behavioral evidence connects directly to real-world AI evaluation beyond benchmark scores.

    Benchmark the full serving system

    Useful measurements include model size, device memory, time to first token, tokens per second, throughput at a latency target, energy per request, loading time, and quality. Compare identical prompts and output lengths after warm-up, then test cold starts and sustained load separately.

    MLPerf Inference defines different serving scenarios with latency, throughput, quality, and compliance rules. Its power methodology measures the full system at the wall for the accompanying workload. This prevents a chip’s theoretical efficiency from being confused with the energy used by memory, cooling, host processors, and software.

    On smaller devices, battery drain, thermal throttling, application size, and competing workloads matter. Our article on on-device AI benchmarking explains why a short, cool demonstration cannot establish sustained mobile performance.

    How to evaluate a quantized model claim

    Ask which weights and activations were quantized, the numeric format and granularity, calibration dataset, supported operators, fallback precision, hardware, runtime, batch size, and sequence lengths. Verify whether the reported speed includes preprocessing, transfers, cache, and output sampling.

    For quality, demand task-specific comparisons against the exact parent checkpoint. A statement such as negligible loss should identify the tests and thresholds used. Also confirm licensing, provenance, and whether quantization changed only representation or included fine-tuning.

    What to watch next

    Lower-bit floating-point formats, better activation handling, mixed precision selected per layer, quantization-aware training, and portable packed layouts will keep improving. Compiler support and fused kernels will decide whether those algorithms translate into real products.

    Quantization is valuable because it turns precision into a controllable engineering resource. The winning deployment is not the smallest model file; it is the system that meets quality, latency, memory, and energy requirements together.

    Primary and authoritative sources

  • Retrieval-Augmented Generation Is a Search System, Not a Truth Engine

    Retrieval-Augmented Generation Is a Search System, Not a Truth Engine

    Retrieval-augmented generation, usually shortened to RAG, is often described as a cure for an AI model’s unreliable memory. The basic idea is sound: search a controlled collection of documents, place the most relevant passages in the model’s context, and ask it to answer from that evidence. The result can be more current, more specific, and easier to inspect than an answer produced from model parameters alone.

    But RAG is not a truth engine. It is a search system connected to a generative system, and each stage can fail differently. A useful RAG product therefore needs careful document governance, retrieval engineering, access control, citations, and evaluation. Simply adding a vector database does not make every answer factual.

    What retrieval adds to a language model

    The NIST definition of RAG describes a generative AI model paired with a separate information-retrieval system or knowledge base. A query triggers retrieval, and selected information is supplied in context for the model to use. This changes the model’s available working material without retraining it.

    The influential 2020 NeurIPS RAG paper combined a model’s parametric memory with a searchable non-parametric memory. The authors highlighted knowledge-intensive tasks, provenance, and the difficulty of updating knowledge stored only in model weights. Modern implementations extend that pattern to company files, product manuals, research collections, code, databases, and web content.

    RAG can make an answer easier to update because the knowledge base can change independently of the model. It can also narrow the model’s attention to approved material. Those advantages are architectural, however; they depend on the quality of the corpus and the retrieval path.

    The pipeline begins long before a user asks a question

    Documents first need to be collected, parsed, cleaned, divided into useful chunks, enriched with metadata, and indexed. Some systems create vector embeddings for semantic similarity. Others combine vector retrieval with keyword search, filters, or a reranker that reorders initial results. Tables, scanned PDFs, diagrams, footnotes, and version histories need special handling because plain text extraction can destroy their meaning.

    Chunk size is a real design decision. Tiny chunks can lose context, while large chunks can bury the relevant sentence and consume the model’s context window. Metadata such as author, publication date, product version, jurisdiction, document type, and access group helps retrieval distinguish two passages that contain similar words but apply to different situations.

    Freshness also needs an explicit policy. A knowledge base can contain a current manual beside an obsolete one, and semantic similarity alone does not know which is authoritative. Ingestion jobs should detect changes, retire superseded material, preserve necessary history, and report indexing failures rather than silently leaving the corpus stale.

    Retrieval can return convincing but wrong evidence

    A retriever may miss the best document because the query uses an unfamiliar term, asks several questions at once, or depends on context from an earlier turn. It may return a passage that shares vocabulary but answers a different question. It may also retrieve several individually relevant fragments that become misleading when combined.

    Query rewriting, hybrid search, metadata filters, reranking, and multi-step retrieval can help, but they add components that must be tested. More retrieved passages are not automatically better. Extra context can distract the model, increase cost and latency, or introduce contradictory evidence.

    This is why the retrieval layer should be evaluated separately from the final answer. Our guide to real-world AI evaluation explains why a single benchmark score rarely describes behavior across users and operating conditions.

    Grounded does not necessarily mean correct

    A response is grounded when its claims follow from the supplied context. That is narrower than factual correctness. The source itself may be outdated, incomplete, biased, or wrong. The model can also summarize a passage inaccurately, confuse two entities, omit an important exception, or answer beyond the retrieved evidence.

    Citations improve inspectability only when they point to the exact material supporting each claim. A link to a long document is weak evidence if readers cannot locate the relevant section. Some systems attach citations after generation, which can create references that look authoritative without actually supporting the sentence.

    The Microsoft RAG evaluation guidance separates document retrieval, context relevance, groundedness, response relevance, and completeness. That separation is useful because a fluent final answer can hide a retrieval failure, while an excellent passage can still be misused by the generator.

    Permissions must survive the retrieval path

    Enterprise RAG often searches material with different access rules. Permission checks should be applied before protected passages reach the model, not merely when a user opens the original file. Indexes, caches, logs, evaluation datasets, and generated answers can all become alternate paths to sensitive information.

    Document content is also an input to an AI system. A malicious or compromised page can contain instructions intended to override the user’s request, reveal data, or influence connected tools. Retrieval filters, content isolation, provenance, least-privilege tool access, and output monitoring are therefore security controls, not optional polish.

    This becomes more important when retrieval is connected to autonomous actions. As our analysis of AI agent protocols notes, interoperability does not provide trust by itself. A system that can send messages or change records needs stronger authorization than one that only drafts an answer.

    A useful test set follows real work

    Teams should build evaluation questions from actual user tasks, including common requests, rare terminology, ambiguous phrasing, multi-turn follow-ups, conflicting documents, missing answers, and permission boundaries. Each case should identify the expected source or acceptable answer when ground truth is available.

    Retrieval metrics can examine whether relevant documents appear near the top and whether irrelevant material crowds them out. Generation checks should cover groundedness, completeness, citation accuracy, correct abstention, and unsupported claims. Operational tests should add latency, cost, freshness, failure recovery, and behavior when the index or model is unavailable.

    Human review remains important for high-impact domains. Automated model judges can accelerate testing, but their scores need calibration against expert decisions. The NIST Generative AI Profile places measurement inside a broader risk-management process that includes governance, documentation, monitoring, and incident response.

    When RAG is useful and when it is unnecessary

    RAG is a strong fit when answers depend on a changing, bounded collection: support manuals, internal policy, regulated guidance, technical literature, or product catalogs. It is less compelling for simple deterministic lookups that a database query can answer exactly. It is also a poor substitute for a calculation engine, transaction system, or professional judgment.

    Users should be able to see which sources were used, open them, and understand when the system lacks evidence. Product teams should publish corpus scope, update cadence, permission behavior, evaluation methods, and known limitations. The provenance lessons in our article on AI content credentials apply here too: evidence can show origin and process, but it does not automatically prove a claim true.

    What to watch next

    RAG systems are moving toward richer document parsing, multimodal retrieval, structured database access, stronger reranking, and retrieval that adapts over several steps. The most important progress will be less visible: better permission enforcement, reproducible evaluation, source-level citations, freshness controls, and graceful refusal when evidence is weak.

    The practical question is not whether a product uses RAG. It is whether the right evidence reaches the model, whether the answer stays within that evidence, and whether users can verify the result. Search and generation can reinforce each other, but neither removes the need for judgment.

    Primary and authoritative sources

  • Confidential AI Protects Data in Use, Not the Entire AI Pipeline

    Confidential AI Protects Data in Use, Not the Entire AI Pipeline

    AI systems often process the information an organization is least willing to expose: private prompts, customer records, proprietary documents, model weights, and the intermediate data created during inference. Encryption protects files at rest and network traffic in transit, but conventional computing still decrypts data in memory while a processor uses it. Confidential computing is designed to narrow that gap.

    The idea is to run a workload inside a hardware-protected trusted execution environment, or TEE. Memory isolation and encryption reduce what the host operating system, hypervisor, cloud administrator, or a compromised neighboring workload can inspect. Remote attestation can then provide evidence about the hardware and software state before an owner releases data or cryptographic keys. That is useful for AI, but it is not a blanket guarantee that the model, application, or result is trustworthy.

    What confidential computing adds to ordinary encryption

    Storage encryption protects a model checkpoint or dataset on disk. Transport encryption protects it while it crosses a network. Both protections normally end when the data reaches the machine that will calculate on it. A privileged administrator, malicious hypervisor, memory-scraping tool, or compromised host could potentially target that exposed working state.

    The Confidential Computing Consortium defines the field around protecting data in use through a hardware-based, attested TEE. Instead of assuming that every layer beneath an application is trusted, the design moves an important part of the trust boundary into the processor. Memory associated with the protected guest or enclave is encrypted and isolated from the surrounding host.

    This changes an infrastructure risk, not the behavior of the model. A TEE can make it harder for the cloud operator to read a private prompt or proprietary weight file. It does not determine whether the model hallucinates, follows an unsafe instruction, or returns sensitive information that the application was authorized to provide.

    Attestation is the admission check

    Isolation is only valuable if a data owner can establish what is inside the protected environment. Remote attestation addresses that problem. Hardware produces signed evidence about identity, firmware, configuration, and measured software state. A verifier compares that evidence with an approved policy. If the result matches, a key broker can release the decryption key or credential needed by the workload.

    NVIDIA’s GPU attestation documentation describes verification of GPU hardware and software claims before confidential-computing modes are trusted. For AI systems, this matters because the calculation may cross a CPU, accelerator, memory, driver, interconnect, and virtual machine. Protecting the CPU while leaving the GPU path outside the boundary would create a conspicuous gap.

    Attestation should be read as a specific claim: an identified platform is in a measured state that satisfies a policy. It is not proof that millions of lines of application code contain no vulnerability. It is also not a timeless certificate. Firmware changes, revoked components, expired credentials, and configuration drift mean verification and policy maintenance are continuing operational jobs.

    AI creates two valuable assets to protect

    The first asset is data. Training, fine-tuning, retrieval, and inference may involve medical records, financial data, unpublished research, internal messages, or personal identifiers. A confidential environment can reduce exposure to infrastructure operators when those inputs are processed. It can also support collaboration in which one party supplies data and another supplies a model without either side handing the raw asset to the other in conventional form.

    The second asset is the model itself. Weights, system prompts, adapters, and inference logic can represent substantial intellectual property. A provider may want to run that model on another organization’s hardware without allowing the host to inspect it. The same isolation and conditional key release can protect both the input and the model package while they are active.

    This is different from the resource questions covered in our article on AI server cooling and compute architecture. Confidential AI adds roots of trust, protected I/O, attestation services, key brokers, and policy decisions to the hardware stack. Those components must scale alongside accelerators and networks.

    The CPU-to-GPU boundary is an engineering problem

    Modern AI inference frequently depends on accelerators, so data may move beyond a confidential CPU guest. Confidential GPU systems attempt to extend protection through the accelerator, including device memory and the connection carrying data between CPU and GPU. Cloud services now expose such configurations, and current NVIDIA platforms provide confidential-computing and attestation mechanisms for supported accelerators.

    The exact boundary still matters. Data may enter through an API gateway, be retrieved from a database, pass through preprocessing code, reach an accelerator, and leave through logging or monitoring systems. A protected GPU cannot compensate for an application that writes full prompts to an ordinary log store. Likewise, a secure VM cannot protect a result after an authorized client downloads it.

    Operators therefore need a data-flow map, not just a product label. They must identify where plaintext exists, where keys are released, which components are measured, what telemetry leaves the TEE, and how updates change the attested state.

    What a TEE does not solve

    Confidential computing does not validate training data, detect bias, prevent prompt injection, or enforce the business purpose for which a result is used. Malicious code running inside an approved TEE can still misuse data it is allowed to decrypt. A vulnerable inference service may expose information through its output even if the host cannot read memory. Model extraction, membership inference, and excessive application permissions remain separate risks.

    It also does not remove supply-chain security. The verifier needs reliable reference measurements and certificate status. Workload owners need a controlled build process, signed artifacts, reviewed policies, and a recovery plan when attestation fails. This echoes our analysis of AI agent protocols: a transport or interoperability layer does not automatically create trust in the software using it.

    Side-channel resistance deserves separate scrutiny. Hardware isolation reduces direct memory access, but timing, resource contention, speculative execution, and application-level behavior can reveal information under some conditions. Security claims should name the threat model and supported configuration rather than saying an AI workload is simply “private.”

    NIST is treating confidential AI as a deployable architecture

    In May 2026, NIST published an initial draft of IR 8320E on confidential computing for cloud workloads. The document presents an example approach for protecting data acted upon by AI in cloud infrastructure. Its draft status matters: it is a technical blueprint for evaluation and discussion, not a final universal certification.

    The NIST framing is helpful because it focuses on a system rather than one processor feature. Machine identity, key management, roots of trust, access policy, and validation all have to work together. That is the same operational perspective needed for AI risk management in critical infrastructure, where a model benchmark is only one part of the control environment.

    Performance and operability still count

    Memory encryption, integrity checks, protected device paths, and attestation add work. The overhead depends on hardware generation, workload shape, memory traffic, network design, and whether an accelerator is involved. A small benchmark under ideal conditions cannot describe every model-serving system. Teams need to measure latency, throughput, startup time, capacity, and failure behavior using their own workload.

    Operational questions may be more important than peak overhead. Can a fleet rotate keys without interrupting service? What happens when an attestation service is unavailable? How quickly can a patched image receive an approved measurement? Can incident responders obtain enough telemetry without leaking the data the TEE exists to protect? A design that cannot be updated or diagnosed safely will not remain trustworthy for long.

    What to watch next

    The strongest progress will make confidential AI easier to verify across vendors. Watch for interoperable attestation formats, clear CPU-to-GPU protection boundaries, reproducible reference measurements, policy-controlled key release, and independent testing of supported configurations. Better documentation should also distinguish production support from preview features and state which devices, drivers, firmware, and orchestration layers are inside the claim.

    Confidential computing fills a real gap by reducing exposure while AI data and models are being processed. Its value is precise: it can change which infrastructure operators and compromised host layers must be trusted. It does not make the model correct, the application authorized, or the output harmless. The mature question is not whether an AI service is “confidential,” but exactly which data is protected, from whom, during which steps, and under what verifiable policy.

    Primary and authoritative sources

  • AI Servers Are Turning Cooling Into Part of the Compute Architecture

    AI Servers Are Turning Cooling Into Part of the Compute Architecture

    Artificial intelligence is changing data-center cooling from a facility utility into part of the computing architecture. Dense accelerator systems concentrate enormous electrical power in a small volume, and nearly all of that power eventually becomes heat. If the heat cannot leave the chip, package, server, rack, and building quickly enough, performance must be reduced or equipment can become unreliable.

    Liquid cooling moves heat more effectively than air near a hot component. That does not make it a universal answer, and it does not remove fans, pumps, heat exchangers, or facility engineering. It creates a connected thermal system that has to be designed, monitored, maintained, and upgraded alongside the processors it supports.

    Cooling starts at the chip package

    An AI accelerator produces heat inside tiny regions of silicon and packaging. That heat passes through interface materials into a heat sink or cold plate. It then travels through air or coolant, into rack equipment, and finally to an outdoor heat-rejection system. Every boundary adds thermal resistance.

    Improving one step does not guarantee a cool system. A highly capable cold plate can still underperform if coolant flow is uneven, a heat exchanger is undersized, or the facility water arrives too warm. Thermal design therefore connects semiconductor packaging to mechanical engineering and building operations.

    Air cooling meets limits as rack density rises

    Air is convenient, electrically nonconductive, and supported by decades of standard equipment. Servers use fans to push it across heat sinks, while computer-room systems remove the warm air. At higher power density, however, moving enough air requires larger heat sinks, stronger fans, careful aisle containment, and more space.

    Fan energy and noise increase, and small airflow imbalances can create hot spots. The US Department of Energy’s data-center design guide notes that high-performance computing facilities adopted direct liquid cooling as rack densities climbed beyond levels that traditional room air systems handle comfortably.

    Direct-to-chip cooling brings liquid close to the source

    In direct-to-chip systems, a metal cold plate sits on a processor, accelerator, or other hot component. Coolant flows through fine channels in the plate and carries heat to a rack manifold. Flexible hoses and quick-disconnect fittings let technicians install or replace server trays.

    The coolant normally remains inside a sealed loop, so liquid does not wash across the electronics. Some components may still rely on air, creating a hybrid design. Memory, network cards, power supplies, and voltage regulators all influence how much heat the liquid loop must capture.

    Rear-door and immersion systems solve different layouts

    A rear-door heat exchanger replaces or supplements a rack door with a liquid-fed coil that removes heat from the server exhaust. It can help a facility raise rack density without redesigning every server, but fans still move air through the electronics.

    Immersion cooling submerges compatible equipment in a dielectric fluid. Single-phase systems circulate a liquid that remains liquid; two-phase systems use controlled boiling and condensation. Immersion can remove heat efficiently, but it changes service procedures, component compatibility, fluid management, and hardware form factors. No one method is automatically best for every site.

    The coolant distribution unit is a thermal interface

    A coolant distribution unit, or CDU, links the clean technology cooling loop serving IT equipment to the facility water loop. It controls flow, pressure, and temperature while a heat exchanger transfers energy between the two circuits. Keeping the loops separate helps protect narrow cold-plate channels from contamination and allows different water chemistry.

    Pumps, filters, sensors, expansion capacity, valves, and redundant controls turn the CDU into critical infrastructure. A design must define what happens during power loss, pump failure, blocked flow, or a rapid change in computing load.

    Retrofitting a building can be harder than cooling a server

    A new data center can place pipes, CDUs, drainage, leak detection, and heat rejection around liquid-cooled racks from the beginning. An existing site may lack floor loading capacity, pipe routes, water connections, or enough room for distribution equipment. Shutting down a live hall to add them can be expensive.

    Open Compute Project guidance emphasizes standardized connection practices between advanced cooling equipment and facility water. Stable interfaces matter because server generations change more often than buildings. A facility should be able to accept new racks without rebuilding the entire thermal plant.

    Reliability depends on materials and maintenance

    Liquid systems introduce risks that air-cooled operations may know less well. Mixed metals can corrode, seals can age, hoses can kink, and particles can obstruct microchannels. Coolant chemistry, cleanliness, pressure control, and material compatibility affect long-term performance.

    Leak detection should find a problem early and identify its location. Quick disconnects need repeated-life testing, and service procedures must prevent contamination. Redundant pumps and bypass paths can keep equipment operating through maintenance, but they add cost and control complexity.

    AI workloads make thermal control more dynamic

    Training and inference do not create perfectly steady loads. A synchronized workload can make many accelerators rise or fall in power together. Cooling controls must respond without allowing chip temperatures to oscillate or wasting energy by running pumps and chillers at maximum output continuously.

    Telemetry from processors, cold plates, racks, CDUs, and the facility can coordinate that response. This is another reason the physical infrastructure belongs in the frontier technology stack, not in a separate conversation that begins after the computers are purchased.

    Efficiency needs more than one headline metric

    Power usage effectiveness, or PUE, compares total facility electricity with electricity used by IT equipment. Liquid cooling can reduce fan and chiller work, especially when warmer coolant enables more hours of compressor-free heat rejection. Yet pumps, CDUs, water treatment, and outdoor equipment still consume energy.

    A fair comparison should also include water consumption, useful computing delivered, equipment utilization, climate, and the energy embodied in replacement hardware. Faster accelerators can finish a job sooner while drawing more power at an instant. Our explanation of why memory and power matter in AI chips shows why system efficiency cannot be inferred from processor specifications alone.

    Cooling does not solve the grid connection problem

    Efficient cooling reduces overhead, but it cannot erase the electrical demand of the computing equipment. Large projects still need utility capacity, substations, backup systems, and a schedule for interconnection. Flexible operations may help some workloads respond to grid conditions, but latency-sensitive services cannot always pause.

    Technologies such as dynamic line ratings can improve use of existing transmission, yet new data-center loads also require conventional grid planning. Cooling and power delivery must be modeled together.

    Higher coolant temperatures can make heat reuse easier

    Air leaving a server room is often too cool and diffuse for economical reuse. A liquid loop can capture heat at a higher, more stable temperature. Depending on local conditions, a heat pump may raise it for district heating, industrial processes, or nearby buildings.

    Open interfaces will determine how quickly the market scales

    Cold plates, manifolds, connectors, CDUs, coolant chemistry, facility loops, and control systems may come from different suppliers. Incompatible pressure ranges, materials, fittings, or telemetry can lock a site into one equipment family. Shared specifications make qualification and replacement more practical.

    The Open Compute Project and ASHRAE are coordinating work across IT and facility boundaries. That type of collaboration is important because cooling systems are expected to outlive several generations of accelerator hardware.

    Limitations

    Published rack-density and efficiency figures usually describe specific hardware, climates, temperatures, and facility designs. They should not be transferred directly to another deployment. Some operational data is proprietary, and long-term reliability evidence for newer connectors and fluids remains uneven.

    Liquid cooling also does not certify the sustainability of an AI service. Model design, utilization, electricity source, construction, water conditions, and hardware turnover remain part of the environmental result.

    What to watch next

    Watch for standardized facility-to-CDU connections, serviceable liquid-cooled server designs, validated material compatibility, workload-aware control systems, transparent water and energy reporting, and designs that can accept several hardware generations. Heat-reuse projects should report delivered heat, not only theoretical potential.

    The future of AI computing will not be decided by chips alone. As power density rises, the path that carries heat away becomes part of performance, reliability, and the pace at which new systems can actually be deployed.

    Sources: Open Compute Project Open Systems for AI white paper; US Department of Energy Best Practices Guide for Energy-Efficient Data Center Design; Open Compute Project connection guidance for advanced cooling systems; ARPA-E COOLERCHIPS program.

  • On-Device AI Needs Better Benchmarks Than Peak TOPS

    On-Device AI Needs Better Benchmarks Than Peak TOPS

    Phone and laptop makers increasingly advertise artificial-intelligence performance with a single large number: trillions of operations per second, usually shortened to TOPS. The figure can describe the peak arithmetic capability of a neural-processing unit, but it cannot tell a buyer whether an AI feature will answer quickly, preserve useful model quality, fit in memory, or keep working after the device warms up.

    That gap matters as generative models move from cloud servers onto personal devices. Local inference can reduce network dependence and keep more data on the device, but it also forces a model to share limited memory, power, and cooling with every other application. A credible benchmark must therefore test the complete experience, not just one component’s theoretical ceiling.

    TOPS measures only part of the system

    TOPS is a rate: how many arithmetic operations a processor may execute under specified conditions. It can help engineers compare hardware blocks when the operation type and numerical precision are the same. Marketing comparisons often omit those conditions. An accelerator may quote a higher figure by using lower-precision arithmetic, while another may run a particular model more efficiently through better software, memory access, or operator support.

    Generative AI is a pipeline. The device must load model weights, process the prompt, move data through memory, schedule work across the CPU, GPU, and neural accelerator, and produce tokens. The slowest stage can dominate. Peak compute that waits for memory or falls back to a less efficient processor will not deliver its headline rate to the user.

    Time to first token and generation speed answer different questions

    A useful language-model test separates responsiveness from throughput. Time to first token measures how long the user waits before an answer begins. Token generation speed measures how quickly the rest arrives. A device can perform well on one and poorly on the other, particularly when model loading, prompt length, or caching changes.

    Benchmarks should publish prompt length, output length, model version, quantization, and runtime configuration. Otherwise, a short warm-cache demonstration can be compared with a longer cold-start task as if they were equivalent. The same discipline applies to multimodal features, where image preprocessing or audio capture may add delays outside the neural accelerator.

    Accuracy belongs beside speed

    A faster answer is not better when model compression removes too much capability. Quantization can reduce model size and energy use by representing weights with fewer bits, but the acceptable trade-off depends on the task and model. Apple, for example, documents optimization techniques that balance model quality, size, and speed rather than treating performance as a single number.

    MLCommons added generative-AI tests to MLPerf Mobile v6.0 in 2026, pairing performance measurements with TinyMMLU and IFEval evaluations. That is an important direction: measure whether a model remains useful while measuring how quickly it runs. It also echoes the broader need for real-world AI evaluation beyond polished demos.

    Memory can be the hard limit

    Model weights are only part of the memory requirement. Inference also needs working buffers and a key-value cache that grows with the conversation context. The operating system, graphics stack, and open applications compete for the same physical memory on many consumer devices. A model that fits during a controlled benchmark may trigger app closures or fail under ordinary multitasking.

    Good reports should list peak memory use, model storage size, supported context length, and whether memory pressure changes performance. They should also distinguish a model permanently bundled with an app from one downloaded later. Storage consumption and update size are real costs even when inference is local.

    Heat and battery reveal sustained performance

    Short benchmark bursts favor peak clocks. Real tasks such as summarizing a long recording, generating many images, or maintaining a voice assistant can run for minutes. As temperature rises, a thin device may reduce power to protect the battery and components. Performance after ten minutes can therefore matter more than the first result.

    A fair test should report device temperature, ambient conditions, power draw, battery energy per completed task, and performance over repeated runs. Apple’s own profiling guidance emphasizes tracing the entire model pipeline and comparing work across the CPU, GPU, and Neural Engine. Those measurements are more informative than assuming every operation lands on the advertised accelerator.

    Local processing offers benefits, but privacy is not automatic

    Running a model entirely on a device can remove the need to send prompts or media to a remote server. It can also support features when connectivity is slow or unavailable. Apple’s Core ML documentation explicitly links strict on-device execution with network independence, responsiveness, and keeping personal data local.

    However, an app may still transmit analytics, sync conversation history, call cloud tools, or fall back to a server for difficult prompts. A benchmark can verify offline operation by disabling the network, but a privacy claim also requires a clear data-flow policy. Local inference describes where computation occurs; it does not describe every destination for the user’s data.

    Model and software updates complicate comparisons

    Two devices with identical chips may perform differently after runtime, operating-system, or model updates. A vendor can improve operator fusion or memory allocation without changing hardware. It can also replace a model with a smaller one that runs faster but behaves differently. Reproducible results need exact software builds and model identifiers.

    This is one reason benchmark governance matters. A shared test suite allows results to be compared under disclosed rules, while independent task-level testing can expose features the suite does not cover. The problem resembles the interoperability challenge discussed in AI agent protocols: a common interface helps comparison, but it does not create trust on its own.

    A practical scorecard for buyers

    Consumers do not need to become chip architects. They need results tied to the feature they intend to use. For an offline assistant, look for time to first token, sustained generation speed, supported languages, accuracy, and whether the task really works in airplane mode. For image tools, look for generation time, output quality, resolution, energy use, and repeated-run performance. For transcription, check real-time factor, speaker conditions, and error rate.

    Reviewers should also state the device memory configuration and battery condition. Synthetic training data and compact models can be valuable, but quality must be monitored; the risks of poorly controlled data loops are explained in our guide to synthetic data and model collapse.

    Limitations

    No benchmark represents every prompt, language, accessibility need, or application. Standard models may not match the proprietary model shipped with a phone. Energy readings can vary with screen brightness, radios, and background services. Quality evaluations also simplify human preferences and may not fully capture hallucinations or safety failures.

    The answer is not one universal score. It is a transparent group of measurements run under repeatable conditions, with quality and efficiency reported together.

    What to watch next

    Watch for mobile benchmark suites to add longer contexts, multimodal workloads, energy per task, cold-start measurements, and sustained thermal runs. Vendors should disclose model identity and precision alongside hardware figures. Independent reviewers can then test the feature as shipped, not merely the processor in isolation.

    Peak TOPS will remain a useful engineering specification. It should be the beginning of an on-device AI comparison, not the conclusion.

    Sources: MLCommons: MLPerf Mobile v6.0; MLCommons Mobile Working Group; Apple Core ML documentation; Apple: Analyzing model runtime performance.

  • AI Agent Protocols Are Becoming Infrastructure, but Not a Trust System

    AI Agent Protocols Are Becoming Infrastructure, but Not a Trust System

    AI agents are being asked to do more than answer questions. They may search company systems, call software tools, delegate work to specialized agents, wait for long-running tasks, and return structured results. That makes interoperability a practical problem. An agent built by one vendor may need to use tools exposed by another platform and cooperate with a third agent that has a different model, framework, and internal design.

    Two open protocols are helping define this emerging stack. The Model Context Protocol, or MCP, standardizes how an AI application connects to tools, data, and other context. The Agent2Agent protocol, or A2A, focuses on communication between independent agents. Together they offer useful building blocks, but they do not make an agent trustworthy, accurate, or safe by themselves.

    Why agents need protocols

    Traditional software integrations usually connect known services through documented application programming interfaces. Agent systems add uncertainty. A request may become a multi-step task, the result may arrive later, and an agent may need to explain what it can do before another system decides whether to delegate work to it. If every platform invents a private method for those interactions, organizations face repeated integration work and stronger vendor lock-in.

    A shared protocol gives developers a common envelope for discovery, requests, results, errors, and lifecycle events. It does not standardize the intelligence inside an agent. Instead, it can make the boundary around that intelligence easier to connect, inspect, and replace. This is similar to the way web protocols allow many kinds of servers and browsers to communicate without making every website identical.

    MCP connects an AI application to capabilities

    The official MCP architecture describes a host, clients, and servers. The host is the AI application. It creates a separate client connection for each MCP server, while a server exposes capabilities such as tools, resources, or prompts. A tool might query a database, create a ticket, or call an external service. A resource can provide context such as a file or record.

    This separation matters because the host remains the coordinator. It can decide which servers are available, what information is passed to them, and when a user must approve an action. MCP focuses on context exchange and capability access; it does not dictate how the host uses a language model or how the model reasons about the returned data.

    A2A connects independent agents

    A2A addresses a different boundary. Its specification is designed for independent, potentially opaque agent systems to discover capabilities, exchange messages, manage tasks, and deliver artifacts without sharing their private memory or internal tools. An agent can publish an Agent Card describing its service, supported input and output modes, skills, and security requirements. Another agent can then decide whether it knows how to communicate with that service.

    The task model is especially important. Some work can finish in one response, while other work needs status updates, streaming, cancellation, or a later result. A shared task lifecycle gives systems a more consistent way to handle those cases. The A2A documentation describes MCP and A2A as complementary: MCP is generally for connecting agents to tools and resources, while A2A is for connecting agents to other agents.

    What interoperability could change for users

    For ordinary users, protocols are mostly invisible. Their value appears when software becomes easier to combine. A travel agent could ask a specialized scheduling agent for available times, use a mapping tool through MCP, and return one coordinated plan. A support agent could delegate a technical diagnosis to a product agent without requiring both systems to use the same model provider.

    Open boundaries may also make systems more replaceable. A business could change one specialist agent without rebuilding every connection around it. That possibility is relevant to the operational governance discussed in Critical Infrastructure AI Needs Operational Risk Management, Not Just Model Benchmarks: organizations need to understand not only a model, but also every connected component and the decisions flowing between them.

    Interoperability is not the same as trust

    A protocol can describe how to send a request, but it cannot prove that the receiving agent is competent or honest. Capability descriptions are claims made by a service. They still need identity, authentication, authorization, testing, monitoring, and contractual accountability. A standards-compliant agent can return a poor answer, mishandle sensitive data, or request more access than it needs.

    The same distinction appears in media provenance. As AI Content Provenance Is Useful, but It Is Not a Lie Detector explains, verifiable history does not establish truth. Likewise, a valid protocol exchange establishes a technical conversation, not the quality of the work produced inside it.

    Security has to be designed at every boundary

    Agent connections can reach files, accounts, databases, and external services, so permissions must be narrow. MCP’s architecture keeps separate client connections to servers, which can support clearer isolation, but implementers still need consent controls and secure credential handling. Official MCP security guidance also discusses risks such as token passthrough, server-side request forgery, session hijacking, and compromised local servers.

    A2A similarly relies on standard web security mechanisms rather than inventing a universal trust system. Production services need encrypted transport, authenticated requests, server-side authorization, input validation, rate limits, and careful handling of push notifications. Delegation should not silently transfer all of the calling agent’s authority. Each service should receive only the access required for the specific task.

    The hard problems remain above the protocol

    Even perfectly compatible agents can misunderstand each other. Two services may use the same JSON fields but interpret a business concept differently. They may disagree about when a task is complete, how confidence is represented, or which evidence belongs in a result. Long chains also add latency, cost, and more places for partial failure.

    Evaluation therefore has to include the whole workflow. The article Synthetic Data Can Train AI, but Reality Must Stay in the Loop makes a related point: internally generated signals are not enough. Multi-agent systems need real tasks, adversarial tests, failure recovery, audit logs, and human review at consequential decision points.

    What to watch next

    The next phase will be less about announcing protocols and more about implementation quality. Watch for stable versioning, compatibility tests, portable identity, granular authorization, clear cancellation behavior, and observability that shows which agent or tool produced each part of a result. Independent conformance testing will matter more as vendors make broader interoperability claims.

    A2A and MCP are promising because they define different layers instead of pretending one interface can solve everything. They can reduce integration friction and make agent systems more modular. The sober conclusion is that open communication is only the first layer. Trust still has to be earned through controls, evidence, and reliable operation.

    Sources: A2A Protocol specification; Linux Foundation: Agent2Agent Protocol Project; Model Context Protocol architecture; MCP security best practices.

  • AI Content Provenance Is Useful, but It Is Not a Lie Detector

    AI Content Provenance Is Useful, but It Is Not a Lie Detector

    AI-generated media has made a familiar internet problem harder: people need to know where a piece of content came from, what changed, and whether it should be trusted. Content provenance is one answer. It does not try to make every fake image impossible. It tries to attach a verifiable history to real media so viewers, platforms, and publishers can inspect the chain of custody.

    That distinction matters. Provenance is not a lie detector. It is closer to a tamper-evident label for digital files. When it works, it can show that a photo came from a particular camera, was edited by particular tools, or was generated by a particular system. When it is missing, viewers still need judgment.

    Provenance Is Different From Detection

    AI detection tries to infer whether content was generated or modified by a model. Provenance records try to document what happened to the content as it moved through capture, editing, export, and publication. Detection is probabilistic. Provenance is evidence attached to a workflow.

    The Coalition for Content Provenance and Authenticity publishes the C2PA technical specification, which defines how signed content credentials can be attached to media. The core idea is that creators and tools can add assertions about origin and edits, then sign them so later systems can check whether the record was altered.

    This connects with our earlier discussion of synthetic data and model feedback loops. Both topics depend on knowing where digital material came from. Without provenance, systems and people may treat generated material as if it were ordinary evidence.

    What a Credential Can Record

    A content credential can include information about capture device, software tool, editing actions, timestamps, identity providers, and whether generative AI was used. The exact fields depend on the tool and policy. A responsible system should avoid exposing unnecessary private data while still giving useful context.

    For example, a newsroom might publish an image credential showing that a file came from a staff photographer, was cropped and color adjusted, and was exported through an approved editing tool. A design app might disclose that an image was generated or expanded using AI. A camera might sign capture metadata at the moment a file is created.

    These records are useful only when viewers can inspect them. That means browsers, platforms, operating systems, editing tools, and publishing systems need consistent ways to preserve and display credentials.

    Signing Helps, but It Does Not Prove Truth

    Digital signatures can show that a credential has not been altered after signing and that it came from a key associated with a particular tool or organization. They cannot prove that the underlying scene was true, that the camera pointed at the right event, or that the signer had good editorial judgment.

    A signed false claim is still false. A credential can be accurate about the editing process while the content is misleading in context. Conversely, an unsigned image may be authentic but lack a machine-readable history.

    This is why provenance should be treated as one signal, not the whole trust system. It works best with editorial standards, source verification, platform labeling, and user education.

    Metadata Can Be Stripped or Rewrapped

    Digital media often passes through apps, social platforms, messaging services, compression tools, and screenshots. Some services strip metadata. Others create derivative files. If provenance data disappears, the viewer may see only an ordinary image with no visible history.

    C2PA is designed to support manifests, signatures, and relationships between original and derivative assets, but real-world preservation requires adoption across the pipeline. A photographer, editor, publisher, platform, and viewer may all use different software.

    This is similar to the smart-home problem we covered in Matter interoperability. A standard can be well designed, but the user experience depends on how many products and services actually implement it correctly.

    Privacy and Safety Need Boundaries

    Provenance systems can expose sensitive information if implemented carelessly. Location, device identity, creator identity, timestamps, and edit history may be dangerous for journalists, activists, minors, or ordinary people sharing personal media.

    Good provenance design therefore needs selective disclosure. A user might prove that content came from a trusted capture process without revealing home address, exact identity, or unnecessary device details. Publishers should understand what they are attaching before they make credentials public.

    Security also matters. Signing keys need protection. If a trusted signing key is stolen, attackers could create convincing false credentials. Revocation and key management are not glamorous, but they decide whether the trust model holds.

    What Platforms Should Do

    Platforms should preserve credentials when files are uploaded, clearly display available provenance, and explain what absence of provenance means. They should avoid treating a credential as automatic proof that content is safe or accurate.

    They should also give users simple controls. A viewer should be able to see whether a piece of media has credentials, who signed them, what edits are disclosed, and whether the file has been modified since signing. That information should be understandable without requiring a forensic tool.

    NIST’s AI Risk Management Framework is relevant here because provenance is part of risk management, not a standalone fix. Organizations need policies for disclosure, monitoring, incident response, and user communication.

    What to Watch Next

    Watch adoption by cameras, editing software, phone operating systems, newsrooms, social platforms, and browsers. Also watch whether credentials survive common actions such as resizing, reposting, screenshots, and format conversion.

    The best case is not an internet where every fake disappears. It is an internet where trustworthy media can carry durable evidence about its origin, and where viewers learn to interpret that evidence with healthy skepticism.

    Sources and Further Reading