AI you can
trust in production
I design the evaluation and reliability layer for RAG and LLM systems — and the QA that keeps software correct — so both behave predictably in production, not just in a demo.
RAG Reliability Audits · AI Security Hardening (OWASP LLM Top 10) · QA & test automation · CI gating · GeoAI.
Most AI systems don't fail at the model.
They fail because nobody measures them.
Confident answers with no grounding in your data.
A prompt, model, or code change quietly breaks quality.
Retrieval that quietly degrades as data grows.
Fixed scope. Fixed price. Clear outcome.
Two service lines, one way of working: start with an audit, then harden, automate, and maintain. No open-ended retainers — every engagement has a defined deliverable.
AI Reliability
RAG Reliability Audit
A fixed-scope diagnostic of your RAG/LLM system: where it's fragile, why, and what to fix first.
AI Security Hardening
Ship safely — OWASP LLM Top 10 in practice
- Guardrails & output validation
- Data-leakage prevention
- CI gate for unsafe outputs
LLM Evaluation Pipelines
Automate quality — never ship a silent regression
- LLM-as-Judge pipelines
- Faithfulness / relevancy scoring
- Regression gating in CI (Azure DevOps)
QA & Test Automation
QA Audit
Find your baseline — a one-time diagnostic of your test coverage: where you're exposed, why it matters, and what to fix first.
1 app / module · recommendations, not implementation
Book the auditAutomation Starter
Get moving — a turnkey delivery project
- Playwright framework setup (config, conventions, CI)
- 20–40 automated tests for critical user journeys
- Azure DevOps pipeline integration (PR gating)
- Docs + 2 review sessions + handover workshop
Blazor SPA / .NET · delivery 4–6 weeks
QA Retainer
Keep it healthy — a monthly partnership
- Guaranteed monthly hours (10 / 20 / 40 h)
- Manual testing of new features before release
- Maintenance & extension of automated tests
- Monthly QA report + Slack / Teams access
Min. 3 months · unused hours don't roll over
GeoAI
AI over spatial data, backed by 17 years in geoscience platforms. A niche where both service lines meet — evaluated for reliability and validated against domain reality, not just unit tests.
From "works in a demo" to "behaves predictably"
RAG & LLM evaluation
LLM-as-Judge pipelines; faithfulness / relevancy / hallucination scoring; retrieval metrics; regression gating in CI (Azure DevOps).
RAG reliability
Retrieval quality (hybrid search, reranking), failure modes, and fallback strategies that hold up under real load.
AI security hardening
OWASP LLM Top 10: guardrails, output validation, data-leakage prevention, and a CI gate for unsafe outputs.
QA & test automation
Playwright test frameworks, coverage & risk audits, and CI/CD pipeline integration (PR gating) for Blazor SPA / .NET apps.
GeoAI
AI over spatial data — PostGIS, QGIS, WMS/WFS — validated against domain reality, not just unit tests.
Outcome: from "AI that works in a demo" to "AI that behaves predictably under real conditions."
17 years where wrong outputs had real consequences
For 17 years at SLB (Petrel, DELFI) I built and tested complex geoscience platforms. That shaped how I work: I don't test whether a system runs — I validate whether it's correct, against domain reality.
Today I work as an independent contractor, remote across the EU, based in Bratislava — under TGS Consult s.r.o..
Building software or an AI feature and unsure if it's reliable?
Start with an audit. Book a 15-minute discovery call — we'll scope where your reliability gaps are and whether an audit is worth it.
Prefer LinkedIn?
Book a discovery call on LinkedIn →Bratislava, Slovakia · Remote across the EU