Independent · Remote across the EU

AI you can
trust in production

I design the evaluation and reliability layer for RAG and LLM systems — and the QA that keeps software correct — so both behave predictably in production, not just in a demo.

RAG Reliability Audits · AI Security Hardening (OWASP LLM Top 10) · QA & test automation · CI gating · GeoAI.

Vladimir Fejdi

Most AI systems don't fail at the model.
They fail because nobody measures them.

→ hallucinations

Confident answers with no grounding in your data.

→ silent regressions

A prompt, model, or code change quietly breaks quality.

→ retrieval decay

Retrieval that quietly degrades as data grows.

Service packages

Fixed scope. Fixed price. Clear outcome.

Two service lines, one way of working: start with an audit, then harden, automate, and maintain. No open-ended retainers — every engagement has a defined deliverable.

AI Reliability

for RAG & LLM systems
HARDENING

AI Security Hardening

Ship safely — OWASP LLM Top 10 in practice

What you get
  • Guardrails & output validation
  • Data-leakage prevention
  • CI gate for unsafe outputs
EVALUATION

LLM Evaluation Pipelines

Automate quality — never ship a silent regression

What you get
  • LLM-as-Judge pipelines
  • Faithfulness / relevancy scoring
  • Regression gating in CI (Azure DevOps)

QA & Test Automation

audit → starter → retainer
RETAINER

QA Retainer

Keep it healthy — a monthly partnership

What you get
  • Guaranteed monthly hours (10 / 20 / 40 h)
  • Manual testing of new features before release
  • Maintenance & extension of automated tests
  • Monthly QA report + Slack / Teams access

Min. 3 months · unused hours don't roll over

Special package

GeoAI

AI over spatial data, backed by 17 years in geoscience platforms. A niche where both service lines meet — evaluated for reliability and validated against domain reality, not just unit tests.

PostGIS QGIS WMS / WFS Spatial data pipelines
ENGAGEMENT
Scoped to your data
PRICING
Contact for pricing
Discuss a project
What I do

From "works in a demo" to "behaves predictably"

RAG & LLM evaluation

LLM-as-Judge pipelines; faithfulness / relevancy / hallucination scoring; retrieval metrics; regression gating in CI (Azure DevOps).

RAG reliability

Retrieval quality (hybrid search, reranking), failure modes, and fallback strategies that hold up under real load.

AI security hardening

OWASP LLM Top 10: guardrails, output validation, data-leakage prevention, and a CI gate for unsafe outputs.

QA & test automation

Playwright test frameworks, coverage & risk audits, and CI/CD pipeline integration (PR gating) for Blazor SPA / .NET apps.

GeoAI

AI over spatial data — PostGIS, QGIS, WMS/WFS — validated against domain reality, not just unit tests.

Outcome: from "AI that works in a demo" to "AI that behaves predictably under real conditions."

Why me

17 years where wrong outputs had real consequences

For 17 years at SLB (Petrel, DELFI) I built and tested complex geoscience platforms. That shaped how I work: I don't test whether a system runs — I validate whether it's correct, against domain reality.

Today I work as an independent contractor, remote across the EU, based in Bratislava — under TGS Consult s.r.o..

17 yrs
QA at SLB — Petrel & DELFI
EU
Remote, independent contractor
OWASP
LLM Top 10 hardening
CI/CD
Eval & test gates in Azure DevOps

Building software or an AI feature and unsure if it's reliable?

Start with an audit. Book a 15-minute discovery call — we'll scope where your reliability gaps are and whether an audit is worth it.

Bratislava, Slovakia · Remote across the EU