Collected sources and patterns will appear here. Add from search or the patterns library.
Image + Prompt -> GroundedMultimodalResponse
Augment vision-language inference by dynamically triggering external search queries based on combined visual and textual input features.
Problem it solves
Vision-language models lack up-to-date or highly specific external factual knowledge when interpreting complex visual inputs.
Consumes
Emits
The real projects this mechanism was found in. Attribution is the point — this is how the best teams actually do it.