OpenAI
GPT-5.x · GPT-5.4 mini/nano
Structured outputs, vision, and reasoning lanes with provider-native wire formats.
AI models
Search engines and procurement teams ask the same question: which models does Vevesh support? Answer: the frontier stack your admins already configure — OpenAI, Anthropic, Google, Groq, Moonshot, DeepSeek, xAI, Qwen, and edge gateways — orchestrated for folder-native RAG, not consumer upload chat.
AI model ecosystem
Vevesh is not locked to one vendor. Our allocator routes engineering RAG, extraction, verification, and massive tool-calling loops across OpenAI, Anthropic Claude, Google Gemini, Groq, Moonshot AI Kimi, DeepSeek, xAI Grok, Alibaba Qwen, MiniMax, GLM, and Cloudflare Workers AI — with metered lanes and provider-native wire formats, not a single-model chat wrapper.
OpenAI
GPT-5.x · GPT-5.4 mini/nano
Structured outputs, vision, and reasoning lanes with provider-native wire formats.
Anthropic
Claude Opus · Sonnet · Haiku · Fable
Anthropic-native tool loops and long-context engineering RAG orchestration.
Gemini 3.1 Pro · 2.5 Pro · Flash
Gemini multimodal extraction, summarization, and Cloudflare edge inference paths.
Groq
GPT-OSS 120B · Whisper STT
Groq-accelerated speech-to-text and ultra-low-latency inference lanes.
Moonshot AI
Kimi K2.5 · K2.6 · K2.7 Code
Massive tool-calling range aligned with Moonshot Kimi code and reasoning models.
DeepSeek
V4 Pro · V4 Flash · R1
Deep reasoning and cost-efficient technical synthesis routes.
xAI
Grok 4.x · Grok Build
Grok-family models via direct and Cloudflare Workers AI gateway lanes.
Alibaba Qwen
Qwen3.5 · Qwen3.7 Plus · Qwen3 VL
Qwen vision-language and plus-tier models for multilingual engineering corpora.
Cloudflare Workers AI
OpenAI · Anthropic · Google · xAI · Qwen
Verified edge gateway ecosystem — route without co-mingling customer KB data.
Together AI
Moonshot Kimi · Qwen · DeepSeek · MiniMax
Open-model hosting for Moonshot, Qwen, and DeepSeek fallback lanes.
DeepInfra
Kimi · Qwen VL · DeepSeek · GLM · MiniMax
High-throughput open-weight routes with metered billing transparency.
Fireworks AI
Kimi · DeepSeek · Qwen · GLM · MiniMax
Premium open-model lanes for vision and code-heavy extraction workflows.
MiniMax
MiniMax M3
Alternative reasoning lane for long-form report generation.
Zhipu GLM
GLM-5.2
GLM-family routes for multilingual standards and annex parsing.
A single-model SaaS cannot optimize every task. Vevesh separates extraction, retrieval, reasoning, and verification into lanes — routing vision-heavy PDFs to Gemini or Qwen VL, fast STT to Groq Whisper, long reports to Claude or GPT, and code-adjacent synthesis to Moonshot Kimi K2.7 Code.
Cloudflare Workers AI provides a verified gateway ecosystem for OpenAI, Anthropic, Google, xAI, and Alibaba models at the edge — while customer knowledge bases stay in isolated data planes. External provider traffic follows explicit privacy tiers; your indexed corpora are never pooled for vendor training.
Supported provider families include: OpenAI (GPT-5.x · GPT-5.4 mini/nano); Anthropic (Claude Opus · Sonnet · Haiku · Fable); Google (Gemini 3.1 Pro · 2.5 Pro · Flash); Groq (GPT-OSS 120B · Whisper STT); Moonshot AI (Kimi K2.5 · K2.6 · K2.7 Code); DeepSeek (V4 Pro · V4 Flash · R1); xAI (Grok 4.x · Grok Build); Alibaba Qwen (Qwen3.5 · Qwen3.7 Plus · Qwen3 VL); Cloudflare Workers AI (OpenAI · Anthropic · Google · xAI · Qwen); Together AI (Moonshot Kimi · Qwen · DeepSeek · MiniMax); DeepInfra (Kimi · Qwen VL · DeepSeek · GLM · MiniMax); Fireworks AI (Kimi · DeepSeek · Qwen · GLM · MiniMax); MiniMax (MiniMax M3); Zhipu GLM (GLM-5.2).
Cursor AI — Complements in-IDE coding — Vevesh owns folder-native RAG and audit citations.
GitHub Copilot — Code completion inside repos; Vevesh indexes specs, reports, and standards beside them.
Microsoft Copilot — M365-embedded Q&A; Vevesh covers mixed engineering folders Copilot does not index natively.
Perplexity — Web search answers; Vevesh grounds on your project files with page-level evidence.
Claude.ai — General chat; Vevesh adds persistent KB indexes, RBAC, and verification loops.
ChatGPT — Session uploads; Vevesh replaces one-off attachments with living folder indexes.
engineering RAGfolder-native AIknowledge base AIenterprise AI knowledge baseOpenAI engineering RAGAnthropic Claude engineeringGoogle Gemini RAGGroq inference AIMoonshot AI KimiDeepSeek RAGxAI Grok engineeringCursor AI workflowChatGPT alternative engineerscitation grounding AIIEC standards AI searchlocal RAG desktopself-hosted AI knowledge basemulti-provider AI gatewayLangGraph tool callingCloudflare Workers AI RAGhybrid vector BM25 searchteam knowledge base RBACzero retention AI privacy
Vevesh is in active development with early design partners. Reach out to join the preview or discuss deployment for your team.
No spam. No sales deck unless you ask. We onboard engineering teams who have real folders to index.