CC-RAGOS turns your real documentation — PDFs, tables, images, help articles — into grounded, cited AI support[2][2] Docling-powered ingestion parses documents other pipelines choke on: scanned PDFs, complex tables, embedded figures. that lives on your own infrastructure and answers where your users already are.
Docling ingests the docs you already have — PDFs, tables, images, help articles — straight into a searchable, cited corpus.
Chat answers with receipts, duplicate-ticket detection, ticket creation, AI-drafted replies — the whole loop, not just a chatbot.
LLM-as-judge evals, satisfaction tracking, agent scorecards, weekly digests — every claim on this page is measured, not asserted.
Questions the docs can't answer become a clustered, prioritized writing backlog — the system gets smarter as it fails.
Not projections — the numbers below are computed live from real platform usage every time this page loads.
The Ask-AI widget and docs-chat answer routine questions before they ever reach the queue — support hours saved around the clock.
Similar-tickets lookup, AI drafts, and tone refinement inside the Support Desk cut per-ticket handle time.
Failed questions become a prioritized docs backlog — fixing the root cause of tickets, not just the symptom.
The weekly report and the MCP server answer management's questions from live data — no manual assembly.
Self-hosted or free-tier cloud (Vercel + Qdrant Cloud + KV) — instead of Intercom/Zendesk-style per-agent pricing.
Deflected conversations per week, handle time before/after, docs written from the gaps panel — all surfaced by the platform itself.
Drag the sliders — prefilled with this platform's real live figures. Estimates, clearly labeled as such.
savings = deflected ticket hours × hourly cost + seats × $39/mo, minus ~$240/yr platform cost. Illustrative only — the section above is the measured reality.
One pipeline, fully owned: no Dify, no black-box vendor. A lightweight FastAPI orchestration layer runs retrieval and ingestion; Next.js renders the evidence.
Docling parses PDFs, tables, images & web docs — formats other pipelines drop.
Agentic chunking splits by meaning, not by character count.
Vectors land in Qdrant — one cloud source of truth, hybrid-searchable.
Hybrid search + metadata filters + rerank pull the right evidence.
Agentic RAG chat composes the reply — every claim pinned to a chunk.
LLM-as-judge evals score faithfulness on real conversations.
Click through — every screen below is the live system with real data, not a mockup.
Full walkthroughs of the deployed surfaces — real data, real answers, no editing. Fullscreen for detail.
Live search, category filters, portal sync, a grounded Ask-AI answer, then the full support-desk dashboard.
The same platform re-skinned for a second client — search, filters, Ask AI and the team dashboard.
A streamed, cited answer with code and screenshots, then the admin view: every chat, feedback & gaps, weekly digest.
Hybrid retrieval over your real docs; the answer cites the exact passages it used. Users can check the receipts — trust is inspectable, not assumed.
Grounded reply drafts, tone chips, similar-past-tickets — the human stays in charge, the AI does the typing. Handle time drops without losing judgment.
Latency histograms, SLA aging, backlog-by-age, staffing heatmaps — every chart clickable to filter the queue. Managers stop asking for reports; the desk is the report.
Every question the docs couldn't answer is logged, clustered by theme, and turned into suggested article titles. The knowledge base grows exactly where users need it.
Our portals reuse Answer's login and Q&A data as the foundation — then add everything its built-in AI can't do.[3][3] Apache Answer's AI assistant (recent releases) only searches content already posted inside Answer — it has no document-ingestion pipeline.
| Capability | Apache Answer | CC-RAGOS |
|---|---|---|
| Knowledge source | Only Q&A posted inside Answer | Any document — PDFs, tables, images, web docs + the Q&A |
| Answer grounding | Black-box AI assistant | Chunk-level citations, visible provenance |
| Quality measurement | No way to measure or tune | LLM-as-judge eval harness on real conversations |
| Where users get help | Destination site they must visit | Embeddable widget + white-label portals on the client's own site |
| Support-ops insight | Votes & reputation | Content-gap detection, feedback panel, weekly digest to management |
| Custom reporting | MCP over its own Q&A | MCP server over cross-portal support + docs-chat intelligence |
| Pipeline control | Fixed, take-it-or-leave-it AI | Own retriever, chunking, model routing — tunable per client |
In one line: Answer stores the knowledge people already asked about. CC-RAGOS turns all client knowledge into cited, measurable, embeddable AI support — and reports back what knowledge is still missing.
A self-hosted, explainable multimodal RAG platform: one FastAPI + Next.js + Qdrant engine powering grounded doc-chat, embeddable help centers, and a full support desk — deployed today for two client portals.
Every answer cites the exact source chunks it was grounded on, and an LLM-as-judge eval harness scores faithfulness on real conversations. If it can't cite it, it says so — and logs the gap.
No — it sits on top. Answer stays the Q&A system of record and login provider; CC-RAGOS adds document ingestion, cited AI answers, deflection widgets, analytics and reporting that Answer doesn't have.
On infrastructure we control: vectors in Qdrant Cloud, app data in Vercel KV, heavy ingestion on our own hardware. The public stats endpoint serves counts only — no user content ever leaves the platform.
Free-tier cloud (Vercel + Qdrant + KV) plus roughly $20/month of LLM API usage. No per-agent seat fees — that alone avoids ~$374/month versus a Zendesk-style desk at current team size.
Model-agnostic via OpenRouter — models are swapped server-side without redeploying, and different tasks (answering, drafting, judging, clustering) can use different models.
They're computed live from platform data on every page load: real conversation and ticket counts, with a clearly-footnoted model for time and cost (15 min per AI answer, 30 min per manual ticket, $20/h, $39/seat — all configurable).
Yes — and before they do, the widget searches past tickets and docs for duplicates, shows matching threads inline, and only then offers "create ticket anyway," pre-tagged and pre-filled.
Heavy ML stays local where the RAM is; serving runs on free-tier cloud. Ingest once locally, query from anywhere.
Four live deployments, one engine. Open any of them — the data on this page came from there.