
What Is a Self-Hosted AI Starter Kit?
A self-hosted AI starter kit packages the core services you need to run private AI on infrastructure you control. Instead of sending every prompt and document to a public AI provider, the full stack runs in your own environment: workflow engine, model runtime, vector database, and application database.
That distinction matters for any business handling ERP data, SOPs, customer records, legal documents, or healthcare information. Sending that content to a third-party API creates exposure you can't fully control.
This article walks through the architecture behind n8n's Self-hosted AI Starter Kit, the Docker deployment path, a sample workflow, and the production considerations most teams underestimate.
Key Takeaways
- Treat the n8n Self-hosted AI Starter Kit as a proof-of-concept starting point, not a production platform
- Stack roles stay clear: n8n runs workflows, Ollama serves models, Qdrant handles vector search, PostgreSQL stores state
- Plan hardware, access controls, backups, and monitoring before treating any deployment as useful or secure
- Validate a simple proof of concept before investing in larger GPU, cloud, or Kubernetes infrastructure
What the Self-Hosted AI Starter Kit Includes
The n8n Self-hosted AI Starter Kit is a Docker Compose template maintained by n8n.
Its own README describes it as a proof-of-concept environment for testing — not something you deploy to production without hardening first. Always check the current repository for component versions and supported profiles before you build on it.
The Four Core Components
| Component | Role |
|---|---|
| n8n | Orchestration layer: receives triggers, moves data between services, calls the model, delivers results |
| Ollama | Local model runtime: downloads and serves models in your environment (runtime and model are separate choices) |
| Qdrant | Vector database: stores embeddings for semantic search and retrieval-augmented generation |
| PostgreSQL | Persistent storage: workflow state, chat memory, and application data (separate from temporary container storage) |

A typical data flow looks like this:
- A user or business event triggers n8n
- n8n retrieves relevant content from Qdrant
- Ollama generates a response using that context
- PostgreSQL or an integrated system stores the resulting state

What It Does NOT Give You
The kit doesn't automatically provide:
- Enterprise identity management
- Compliance validation
- High availability
- Comprehensive observability
- Disaster recovery
- Guaranteed answer accuracy
It's a starting architecture. Everything past "it runs" is on you.
Plan Requirements: Hardware, Models, and Docker
Start with the business problem, not the model. Define your documents, users, response-time expectations, privacy requirements, and how much human review each answer needs before you pick hardware.
Hardware Trade-offs
There's no universal RAM or GPU minimum. Ollama runs CPU-only or on NVIDIA, AMD ROCm, and Apple Metal. Qdrant sizing depends on vector count, dimensions, payload indexes, storage, and quantization—not a fixed RAM number.
Broadly:
- CPU-only — fine for smaller models and light usage; expect slower responses
- Consumer/workstation GPU — better for moderate concurrency and faster inference
- Dedicated server GPU — for larger models or higher concurrency
- Private cloud — trades capex for ongoing runtime cost while keeping data off public AI endpoints
Choosing a Model
Weigh these factors, not just which model is trendy:
- Accuracy for your specific task
- Context-window size needed
- Language and tool-use support
- Latency requirements
- Hardware compatibility
- Licensing terms
- Sensitivity of the data it will process
A smaller local model is often enough for classification, extraction, summarization, or internal search. Complex multi-step reasoning may need a more capable model. Test on your real documents and workflows—don't assume bigger is better.
Docker Prerequisites
Before deployment, plan for:
- Container CPU/memory limits and reservations
- Persistent volumes for each service
- Environment variables and secrets (never hardcoded passwords)
- Service-to-service networking by name, not exposed host ports
- Which profile matches your hardware — CPU, NVIDIA GPU, or AMD GPU
Check current official documentation for supported profiles rather than relying on commands from an old tutorial. Compose files change.
The stack is free and open source; running it is not. Budget for:
- Infrastructure (on-prem or private cloud)
- Storage growth
- Electricity or cloud runtime
- Updates and ongoing administration
Deploy the Self-Hosted AI Starter Kit Step by Step
Work through these steps in order. Confirm each layer before you move on.
- Clone the official repository, review the README and
.env.example, and record the exact version you're testing - Choose the matching Docker Compose profile for your hardware (CPU, NVIDIA GPU, AMD GPU), then start services and check container logs if something fails
- Access n8n for the first time, create the owner account, configure credentials, and confirm a model is available in Ollama
- Verify each layer independently:
- Can n8n reach Ollama?
- Does the model answer a basic prompt?
- Can Qdrant store and retrieve a test embedding?
- Does PostgreSQL data survive a container restart?
- Test with non-sensitive sample data first and document every configuration change you make

Troubleshooting Checklist
If a service fails to start or respond, check these first:
- Model not pulled or unavailable in Ollama
- Wrong service hostnames (use Compose service names, not
localhost) - Ports exposed publicly by accident
- Not enough memory for the selected model
- Missing volume mounts or permission errors on mounted directories (data loss on restart)
- Containers starting before dependencies are ready
Don't expose the n8n interface publicly until authentication and network controls are in place. This is where a lot of quick demos turn into security incidents.
Build a First Private AI Workflow
Document Q&A is the fastest private AI workflow to stand up—and the one most teams can validate with real internal content in days, not months.
- n8n receives an approved file from a controlled source
- The content is split into chunks and converted into embeddings
- Those vectors are stored in Qdrant for retrieval
- On each query, only the relevant passages are retrieved
- Only that retrieved context—not the whole document—is sent to the local model
- n8n returns the answer with source references attached

Common business uses include:
- An ERP onboarding assistant trained on approved documentation and SOPs
- A private support assistant for manufacturers or distributors answering questions about internal procedures, inventory, or orders
- An internal search tool for operational knowledge that shouldn't leave the building
Querying a Database Safely
Once document Q&A is working, many teams want the same natural-language pattern over live business data. If users can query a database, do not give the model unrestricted access:
- Use a restricted, read-only connection
- Map user roles to specific permitted tables or fields
- Validate generated queries before execution
- Never allow write-back or schema changes
This is the same pattern AI-ABW uses for ERP and SQL Server integration: authorized employees get natural-language access to approved data through controlled, read-only connections. The system reports on the data; it does not change records in the source system.
Before rolling out to real users, evaluate against:
- Representative test questions and expected answers
- Retrieval quality checks
- Hallucination spot-checks
- Latency observations
- Human approval required for any high-impact decision
Self-hosting alone does not prove HIPAA compliance or preserve legal privilege. Confidential or regulated data still needs a formal privacy and compliance review before production use.
Secure and Move from Demo to Production
A working demo and a production system are not the same thing. Here's the gap you need to close.
Baseline Security
Treat the starter kit like an internal system from day one:
- Private networking, not public exposure
- Strong authentication on every service
- Least-privilege service accounts
- Restricted inbound ports
- Encrypted connections where appropriate
- Protected secrets (not stored in plain text in Compose files)
- Timely patching and logged administrative access
Data Governance
Security controls only hold if you also decide what data the stack is allowed to keep. Decide upfront: which prompts, documents, embeddings, chat histories, model files, and logs get retained? Define retention, deletion, backup, and access policies before connecting live business data — not after.
Backups That Actually Restore
Production also means you can recover when something fails. Back up PostgreSQL, workflow definitions, configuration, and vector data. Then test the restoration. A mounted volume is not a backup plan; it's a single point of failure if the host disk dies.
Monitoring
Once the system is live, you need visibility into failure modes before users do. Track container health, model response failures, latency, resource usage, disk growth, failed workflows, and unexpected access attempts. Set up alerting, and redact logs where sensitive prompts might appear in plain text.
Demo vs. Production
| Starter Kit POC | Production Architecture |
|---|---|
| Single Compose file | Reverse proxy, identity integration |
| Manual scaling | Workload queues, GPU scheduling |
| No staging | Staging environment, version control |
| Basic logs | Full observability, incident response plan |

Re-evaluate performance after any change to the model, prompt, embedding strategy, or retrieval logic. A small update can quietly shift answer quality or how data gets handled.
If you want a privately hosted business AI assistant on ERP documentation, SOPs, or controlled business data — without building and maintaining this stack yourself — AI-ABW provides a supported private path grounded in 30+ years of enterprise software work, with company data kept off public AI systems.
Frequently Asked Questions
How do you choose an AI model for self-hosting?
The best model depends on your use case, hardware, required accuracy and latency, context needs, and data sensitivity. Test a few candidate models against your actual business tasks rather than picking a "best overall" model.
How good are self-hosted AI models?
Self-hosted models handle summarization, extraction, classification, document search, and internal assistants well. Performance depends heavily on model choice, hardware, retrieval design, and prompting quality.
Can you run AI on Docker?
Yes. Docker can package the model runtime, workflow engine, vector database, and relational database into connected services. It simplifies deployment, but it doesn't remove your hardware, security, or maintenance responsibilities.


