Blog
/
Article

Memories.ai Brings Omni-Model to Snapdragon® to Accelerate On Device Perception

Memories.ai's efficient video-language model stack now runs natively on Qualcomm® Snapdragon® platforms, extending LUCI's on-device personal memory across phone, PC, and smart glasses.

Memories
Memories
·
September 24, 2026

At Snapdragon® Summit 2026, Memories.ai announced that its video-language model stack now runs natively on Qualcomm® Snapdragon® platforms, enabling continuous, on-device visual perception across phone, PC, and smart glasses. The first consumer expression of that stack is LUCI, Memories.ai's personal AI, which indexes, remembers, and retrieves a person's own context on their behalf, with the full perception and retrieval loop running locally.

The compute problem with continuous perception

Video understanding has, until now, been a cloud workload. Frontier vision-language models run in the single-digit billion of parameters, expect batched offline inference, and assume a clip has already been uploaded before anyone asks a question about it. None of those assumptions hold up against a device that is perceiving continuously. A wearable or a PC that observes a person's day produces hours of visual input daily, inside a fixed thermal and power budget, on data most people would rather not send to a server in the first place.

Memories.ai's approach restructures where the compute goes, rather than shrinking a cloud model and hoping it fits.

Architecture: amortized perception, lightweight retrieval

OmniCaptioner is an efficient video-language model that runs at capture time. It converts what a device sees and hears into a compact, structured, temporally grounded description of the moment, the only step that touches raw pixels, and it runs once per moment.

OmniRetriever is a multimodal retrieval model that indexes those descriptions into a queryable personal memory, then resolves natural-language queries against it at interactive latency.

The result is that heavy visual understanding gets paid for once, at capture, instead of being paid again at every query. Everything downstream, indexing, retrieval, and answer synthesis, runs on compact representations rather than on video, which keeps every model in the loop small enough to live on an NPU rather than in the cloud.

Why this maps onto Snapdragon®

Quantized and compiled for the Qualcomm Hexagon NPU, models at this scale fit inside the sustained power envelope of a phone, a laptop, and eventually a pair of glasses. That matters for three reasons:

  • Continuity. Perception can stay on rather than being invoked. Memory becomes a function of what was observed, not of what a person remembered to capture.
  • Privacy by architecture. The index never leaves the device. The cloud, when used at all, connects context across a person's devices, it doesn't store or process the underlying personal content.
  • Cost. A cloud VLM running continuously per user isn't a viable unit economic. An on-device one is roughly free at the margin.

A full day of continuous memory compresses to a couple of gigabytes on-device.

One memory, three surfaces

LUCI shipped first on the desktop, indexing on-screen activity in real time. At Snapdragon Summit, Memories.ai extended the same architecture across every device a person carries:

  • LUCI Mobile, on Snapdragon® 8 Elite: the full pipeline, indexing, retrieval, transcription, OCR, and answer generation, runs on the Hexagon NPU, with the phone doubling as the hub where memories from other surfaces consolidate.
  • LUCI Desktop, on Snapdragon® X2 Elite: continuous screen memory for an AI PC, so everything seen on screen, meetings, documents, browsing, is understood and searchable locally.
  • A real-time visual memory demo on Snapdragon® AR1 Gen 1, running on RayNeo smart glasses, performs live indexing of what the wearer sees with natural-language recall later. This is a technology demo, not yet a shipping product.

A moment captured on one device becomes queryable memory on any of the others, with no centralized cloud index required.

Why personal AI needed this

Most assistants today are stateless. Every session starts from zero context, because the only memory available to them is what a person types into a chat window. LUCI closes that gap by giving the model a persistent, queryable memory built from what a person actually sees, hears, and does.

"The interesting engineering result here is not that we made a model smaller," said Shawn Shen, founder and CEO of Memories.ai. "It is that we moved the expensive part of video understanding to capture time, once, so that everything after it fits on an NPU. The Qualcomm Hexagon NPU is what lets that pipeline run continuously and privately, on the device itself. Personal AI is the first place this pays off, but the same stack is what physical intelligence will need generally."

"At Qualcomm Technologies, we believe the next generation of AI should be personalized, highly capable, and privacy-first," said Vinesh Sukumar, Vice President, Product Management, Qualcomm Technologies, Inc. "Memories.ai's LUCI demonstrates how on-device AI can deliver meaningful experiences that understand context and provide intelligent assistance while keeping personal data under a user's control. By leveraging Snapdragon's AI capabilities across phones, PCs, and smart glasses, Memories.ai is helping unlock a new era of connected, personalized AI experiences that seamlessly enhance user's everyday experiences."

What's next

Memories.ai's approach to visual memory doesn't stop at personal devices. The same architecture, amortize understanding at capture, keep everything downstream small and local, is the foundation the company is building toward for physical intelligence more broadly: any system that needs to understand and remember the physical world it operates in.

For now, LUCI is where it shows up first: a personal AI that sees, remembers, and retrieves, entirely on the device in your pocket.

Read more