Agent Context
Generate agent-facing multimodal scene context—including semantic scene graphs, visible-object states, and Set-of-Mark (SoM) visual snapshots—and query Gemini to inspect and describe objects in the spatial environment.
[!NOTE] For local prototyping, the pre-launch dialog accepts a Gemini API key stored locally in your browser. Production applications must proxy AI requests through a secure server.
Desktop Controls (User Mode)
- Capture & Query Context: Left-click to capture scene context and ask Gemini to describe the scene. Open the browser developer console (F12) to inspect the captured context structure.
- Move Camera: Use W, A, S, D to move forward, left, backward, and right.
- Elevate Camera: Use Q and E to move down and up.
- Rotate: Right-click and drag the mouse to look around.
XR Controls
- Capture & Query Context: Ray-pinch with hands or pull the controller trigger to capture spatial context and query the AI agent.
main.js