Reflections

Proactive AI for smart glasses

2026 · Source on GitHub

Motivations

Voice assistants on glasses have an inherent limitation. Whether it’s Meta’s Live AI or Gemini Live, all have the same shape: a wake word, a question, an answer. Fundamentally, these systems are reactive. They wait for the user to declare intent. The ideal system is proactive. It recognizes intent before the user declares it. That’s what we built. A proactive system, fed by everything smart glasses see and hear.

The system at a glance

The system is composed of three layers: sensory, proactivity, and memory. At a high level, the sensory layer extracts context from the glasses input, the proactivity layer determines whether to intervene, and the memory layer persists information learned about the environment.

Demo

Sensory

We wanted Reflections to be driven by the conversations we have with other people. The sensory layer takes in the audio and video streams collected by the glasses and extracts a live, diarized transcript of the current conversation, with real identities mapped to who says what. On the audio side, we used the Soniox API to do real-time transcription along with speaker based diarization. On the video side, faces are detected via YuNet and tracked frame to frame. LR-ASD scores which of those faces is currently speaking.

Fusion is where this all comes together. Soniox tells the system that speech happened and which voice produced it. ASD tells the system which face in view was the one talking. When ASD is uncertain (no face visible, two faces both scoring positive) the diarization ID still keeps distinct voices distinct in the transcript, and the face to voice mapping stays preliminary until the situation clears.

Identity has its own lifecycle. A new face enters the identity gallery as “Person N”, pending. Once that face is confirmed speaking, the slot is promoted: “Person N” becomes a stable, anonymous participant, and every subsequent line from that voice attaches to it. At the end of the session, Claude Haiku reads the full transcript and proposes real-name mappings for those slots, looking for moments where someone is addressed by name, introduces themselves, or is referred to in the third person.

Individually, none of these models are enough for robust identity recognition. But woven together with fusion logic, Reflections is able to recognize people it has seen before and know exactly who is saying what.

Proactivity

The proactive agent determines when to intervene by repeatedly assessing the current situation using context from the sensory layer as well as memory. Continuous calls to cloud model providers would be prohibitively expensive for practical use, so Reflections uses a two tiered architecture to minimize cost while maximizing responsiveness. A small, fine tuned proactivity classifier is repeatedly called locally, only passing requests to the large LLM if deemed that the situation could require intervention.

On every transcript update, the sensory layer context, relevant memories, and available tools are passed to the classifier, a Qwen 3 1.7B model with a LoRA adapter we trained on curated conversation traces and published to Hugging Face. The model returns one number: P(actionable | context). If that probability clears P ≥ 0.25, the pipeline continues. Below that threshold, Reflections stays silent. The whole gating step runs in about 200 ms on a Mac M-series GPU.

We trained the classifier via LoRA on examples that look like

<label>{0|1}</label><signal>…</signal><reasoning>…</reasoning>

Label first, then the rationale. The model sees full reasoning during fine-tuning, but at inference time we stop the prompt at <label> and just read the next token softmax over 0 and 1. No chain of thought is generated, which significantly lowers response latency. Chen et al. (2024) has shown this technique of rationale-after-label training transfers well to label-only inference.

When the probability does clear the threshold, the transcript gets sent to Claude Haiku, along with contents from the memory layer and a set of tools: web search, Google Maps, Calendar. The system can either stay silent, take an action through a tool, or speak. Only if it chooses to speak does Reflections send anything back to the wearer.

A log of classifier prompts, gate decisions, Claude requests and responses, and memory updates, one line per event with timestamps.
The proactivity dashboard: every classifier pass, gate decision, and model call, in order.

Memory

The agent keeps a single file, memory.md, that grows as a curated gallery of the people in the wearer’s life. One Entities section, one Name (relationship) block per person, and bullets underneath for persistent facts like preferences, work, or habits. After testing multiple sophisticated systems like knowledge graphs and note linking, we found that a markdown file was sufficient for our usage.

The memory is read upon every classifier pass using a regex parse, so that the prompt only sees entities relevant to this conversation. Every distinct speaker in the live window is fuzzy-matched against entity names and aliases and put into a relevant people block Claude sees. As the memory file grows, this selectivity prevents bloat in the proactivity prompt.

Closing thoughts

We present a few key takeaways from building and deploying Reflections in real world environments.

  • Proactivity is only as good as the sensory input you feed in. Errors from facial detection, speaker diarization, active speaker detection, and transcription flow downstream into memory and context inaccuracies.
  • The hardware is not there yet. We initially tried using Meta Ray-Ban glasses, however the video quality via SDK was not sufficient. Mentra glasses’ wifi stream was better, but the quality is still not 1080p level and the chip tends to overheat, preventing full 24/7 use.
  • The potential is huge. With a more robust sensory layer, better hardware, additional tools, and a feedback / personalization system, smart glasses have the potential to function in daily use and act as an extension of the human mind. We were amazed by the utility we were able to uncover in our testing just from the few months we spent building Reflections.