Structured extraction
Qwen3.5 0.8B
WebGPU generation uses a 65,536-token context with F16 KV cache via pinned wllama/llama.cpp. The Rust/Candle CPU fallback remains capped at 2,048 input tokens.
Not loaded
Results appear here.
Semantic retrieval
EmbeddingGemma 300M
This path implements bidirectional Gemma attention, mean pooling, both official dense projections, Matryoshka truncation, and L2 normalization in Candle. Use is subject to the Gemma terms.
Not loaded
Results appear here.