Skip to main content

Agent Context

Generate agent-facing multimodal scene context—including semantic scene graphs, visible-object states, and Set-of-Mark (SoM) visual snapshots—and query Gemini to inspect and describe objects in the spatial environment.

[!NOTE] For local prototyping, the pre-launch dialog accepts a Gemini API key stored locally in your browser. Production applications must proxy AI requests through a secure server.

Desktop Controls (User Mode)​

  • Capture & Query Context: Left-click to capture scene context and ask Gemini to describe the scene. Open the browser developer console (F12) to inspect the captured context structure.
  • Move Camera: Use W, A, S, D to move forward, left, backward, and right.
  • Elevate Camera: Use Q and E to move down and up.
  • Rotate: Right-click and drag the mouse to look around.

XR Controls​

  • Capture & Query Context: Ray-pinch with hands or pull the controller trigger to capture spatial context and query the AI agent.
main.js