Self-hosted · runs on your hardware

Your supportknowledge,answering 24/7.

Cloud5 Keep is a self-hosted AI support assistant for companies that sell technical products. It recommends the right equipment, troubleshoots faults, and reads diagnostic logs — in your customers' language, with zero data leaving your infrastructure.

No per-message AI feesNo data leaves your networkNo vendor lock-in
01 — The problem

Off-the-shelf bots
answer fast — and
leak everything.

Technical support is expensive, slow, and hard to scale. Your best engineers spend hours answering the same product-comparison and troubleshooting questions. Knowledge lives in PDFs and people's heads.

Cloud chatbots solve part of this — but they ship your customers' data and your proprietary documentation to a third-party cloud, bill you per message, and give generic answers that don't know your catalog.

Keep runs on hardware you own, learns only from your documentation, and answers with the precision of someone who has read every spec sheet you've ever published.

02 — What it does

One assistant, four jobs.

Built for technical catalogs — networking, servers, storage, industrial systems — where customers ask detailed questions and expect exact answers.

01
Recommend
Product & equivalent matching
“A 1U switch with 48× 10GbE and PoE+” or “the equivalent of a competitor's model” returns up to three ranked picks from your catalog — with the specs that justify each one.
02
Resolve
Troubleshooting guidance
Draws on your manuals and a growing library of resolved cases: probable cause, step-by-step fix, prevention. When confidence is low, it escalates — no wrong answer reaches the customer.
03
Diagnose
Log & diagnostic analysis
Paste syslog, IPMI/iDRAC/iLO events, or network device logs. Keep detects the format, extracts the signal, and returns a structured diagnosis: what's wrong, root cause, fix, prevention.
04
Reach
Every channel, every language
A website widget and WhatsApp Business sharing one brain. Native English and Simplified Chinese, detected automatically per message. Text, files, voice, and screenshots in.
03 — One deployment

One brain.
Three audiences.

The same knowledge base, the same escalation flow — serving everyone who touches your product. One investment, three returns.

Customer-facing
Your customers
24/7 answers on your website and WhatsApp — product questions, troubleshooting, diagnostics — without waiting on a human.
Workforce
Your support team
A shared, always-current knowledge base plus an agent dashboard to watch, steer, or take over any conversation.
Onboarding
Your new hires
A train-up engine. New staff learn the catalog and the runbooks by asking — instead of reading every manual on day one.
04 — The economics

Why self-hosted
changes the math.

For thousands of conversations a month, the gap between per-message cloud billing and self-hosted inference is the gap between a four-figure monthly bill and effectively zero.

Cloud AI chatbots
Cloud5 Keep
Where your data lives
Third-party servers
Your hardware, always
AI inference cost
Per-message, forever
$0 — runs on your GPU
Customer data exposure
Sent to external API
Never leaves your network
Knowledge ownership
Uploaded to vendor cloud
Stays in your systems
Scaling cost
Rises with every conversation
Flat — hardware is one-time
Recurring fees
Subscription + usage
Messaging fees only*

* The only external recurring cost is WhatsApp Business messaging fees (paid to Meta, not to us), if you enable the WhatsApp channel.

05 — Built for control

Yours, end to end.

Single-tenant by design. No shared infrastructure, no co-mingled data, no surprises.

Single-tenant
Your deployment is yours alone. No shared infrastructure, no co-mingled data.
Your knowledge, your rules
Your team writes and maintains everything the bot knows. Publish an update and it's live to customers in under a minute.
A human, one step away
Low-confidence answers, explicit requests, and critical faults route to your team — with the full transcript, contact, and a reference number attached.
Watch, steer, or take over
Agents can monitor any live conversation, quietly guide the AI's next answer, or take over entirely — without the customer noticing a handoff.
06 — Trust

No wrong answer
reaches the customer.

The single biggest risk with AI support is a confident, wrong answer. Keep is built so that can't happen.

01
Strict grounding
Answers are generated only from your documentation — never from the model's general training.
02
Source verification
A verification gate checks every answer against its retrieved sources before it's allowed to send.
03
Refusal on unknown
Asked about something not in your knowledge base, Keep says so plainly — it does not invent.
04
Auto-escalation
Low confidence or a critical fault hands off to a human automatically, with full context.
07 — How deployment works

From kickoff to
production in 8–12 weeks.

The full stack deploys to your server. Nothing runs in our cloud.

01
Discovery & scoping
We map your catalog, documentation, channels, and escalation workflow.
02
Install on your hardware
The full stack deploys to your server. Nothing runs in our cloud.
03
Knowledge ingestion
Your manuals, spec sheets, and runbooks are loaded and indexed.
04
Tuning & QA
Retrieval quality, language coverage, and escalation logic validated against real queries.
05
Go live
Website widget and WhatsApp channel activated. Your team trained on the dashboard.
06
Ongoing
The knowledge base grows automatically from every resolved case.
08 — Under the hood

Production-shaped by design.

Every tool maps one-to-one to the production stack, so scaling up is a configuration change — not a rewrite. We're honest about what's proven today versus delivered in a production engagement.

Component
Proof of concept
Production
Language model
Qwen2.5-14B · Ollama
Qwen2.5-32B AWQ · vLLM
Vision model
Qwen2-VL-7B (co-resident)
Embeddings
bge-m3 (multilingual)
bge-m3 (multilingual)
Speech-to-text
Whisper large-v3
Re-ranking
bge-reranker (cross-encoder)
bge-reranker (cross-encoder)
Vector store
Qdrant · hybrid
Qdrant · hybrid
Backend API
FastAPI
FastAPI
Database
SQLite
PostgreSQL
POC scope honesty: the proof-of-concept runs on a trusted network without auth or rate limiting — by design. Production hardening (auth, rate limiting, abuse prevention) is part of the production engagement.
FastAPILlamaIndexQdrantbge-m3Qwen2.5vLLMWhisper large-v3Cross-encoder rerankHybrid searchPostgreSQLRedisMinIODockerSelf-hostedOn-premEN + 简体中文FastAPILlamaIndexQdrantbge-m3Qwen2.5vLLMWhisper large-v3Cross-encoder rerankHybrid searchPostgreSQLRedisMinIODockerSelf-hostedOn-premEN + 简体中文
09 — Recommended hardware

What it runs on.

Production tier (32B + co-resident vision) wants a 24 GB+ class GPU. The POC validates the full architecture at a 14B tier on a 12 GB GPU.

GPU
NVIDIA RTX 5090 (32 GB) or 24 GB+ class
CPU
Dual-socket Xeon or comparable
RAM
≥ 64 GB (128 GB comfortable)
Storage
≥ 1 TB NVMe SSD
Network
Static IP / DDNS · ports 80/443 for WhatsApp webhooks
11 — Proven today

Not a pitch. A working build.

The proof-of-concept is a complete vertical slice of the production architecture for product recommendation, running on the intended stack. It already demonstrates:

  • SKU disambiguation and accurate spec retrieval
  • Comparison and superlative reasoning across the catalog
  • Anti-hallucination: source verification, strict grounding, refusal on unknown items
  • Follow-up / anaphora resolution (“tell me more about the one you mentioned”)
  • English and Simplified Chinese
  • Persistent conversations, streaming responses, file ingestion

Because every tool maps one-to-one to the production stack, migration is a configuration change — not a rewrite. The foundation is production-shaped by design.

10 — Engagement

One-time build. You own it.

No per-message AI billing. No software subscription. Ongoing recurring cost is electricity and — only if you enable WhatsApp — Meta's messaging fees.
Full deployment
Production
$0/message
One-time build — you own the stack forever
Full production stack on your hardware
Knowledge ingestion + auto re-index
Web widget + WhatsApp · voice · vision
Escalation + live agent dashboard
English + Simplified Chinese
Team training + 8–12 week delivery
Proof of concept
2–3 weeks
Validate on your real catalog first
Product-recommendation vertical slice
Your real SKUs + spec sheets
Hybrid retrieval + cross-encoder rerank
Anti-hallucination grounding gate
English + Simplified Chinese
Web widget · streaming responses
Not sure which tier?
Log analysis, voice, vision, multi-site, custom channels — scoped to your stack.
Book a 30-minute scoping call. No commitment required.
Book a scoping call →
12 — Request a demo

Have data you can't send to the cloud?

That's exactly who we built Keep for.
Tell us about your catalog and your channels. We respond within 24 hours and offer a free 30-minute scoping call — no commitment required.
hello@cloud5.tech
Dhaka, Bangladesh — deployed globally
Response within 24 hrs
Self-hosted · single-tenant · your hardware
Now booking proof-of-concept builds for technical-product companies.