Support systems & AI engineering

I build support systems that scale before headcount has to.

I work across support operations, data, automation, and LLM systems: turning repeated queue pressure, fragmented knowledge, and brittle manual processes into tooling a team can actually operate.

Da Nang · GMT+7Async across time zonesRegulated, integration-heavy systems
For a bounded operational problem, send the short brief →
SUPPORT-LED SYSTEMSWork starts with real queue, incident, and handoff conditions—not a vendor benchmark.
Illustrated portrait of Oron Culzac
operating pathlive map
SIGNALTickets, logs, alerts, reviews
STOREProduction data and history
DIAGNOSESQL, incidents, root cause
OBSERVEMetrics the team can use
ACTGated automation and handoffs
ASSISTRetrieval and agent-facing AI
L2/L3 support context

Queue, incident, escalation, and handoff work.

SQL-led diagnosis

Operational investigation before automation.

Prototype → production

n8n for speed; code where control proves necessary.

RAG evaluation

Retrieval, grounding, injection, and permissions.

Operability first

Human gates, idempotency, observability, handover.

Selected work

Outcome first. Architecture underneath.

Three systems that show the range: knowledge retrieval, safe automation, and AI that helps an agent without being given the wrong authority.

Knowledge systemsEvaluated

Permission-scoped support knowledge RAG

A pipeline from raw support exports to a deduplicated, citation-bearing corpus, built for L2/L3 triage rather than generic chat.

I built the ingestion, access model, hybrid retrieval, evaluation harness, and embedding infrastructure end to end.

Recall@1 91.8%
389 held-constant cases
Technical notes

Eleven sources normalised into one schema; 74 scripts; canonical knowledge units before embedding; semantic, bilingual full-text, and exact-term retrieval legs. Evaluation also covered grounding, citations, prompt injection, and permission boundaries.

Support automationProduction ownership

From n8n prototype to a production support chatbot

A proof of concept running on real channels before the organisation committed to a production service.

I built the prototype; two platform developers rebuilt it in async Python/FastAPI. I did not write that service. I later took ownership of maintenance and hardening.

78-node prototype
approved for productionisation
Delivery notes

The original prototype handled deduplication, envelope normalisation, Telegram identity resolution, two-way mirroring, and decision paths for answer, escalation, and handoff. It proved the architecture and vocabulary before promotion.

Agent-facing AIARCHITECTURE / PHASED ROLLOUT

Zendesk copilot with asymmetric-risk controls

An agent-facing architecture for drafted replies and durable summaries inside the Zendesk workspace, where people can accept, edit, or ignore them.

The design keeps generation agent-facing, defines deterministic boundaries around side effects, and binds artifacts to prompt-version and model records for evaluation.

2 architectures
12 tracked versions
Architecture notes

The design covers a webhook architecture and a polling variant with Postgres cursor state, lease-based claiming, and reaper recovery. Triage and run-level observability are deliberately deferred rather than claimed as shipped.

Internal figures are scoped project measurements, not independently publicly inspectable. Status, contribution boundaries, and definitions matter as much as the numbers.

More operational systems

ROLLOUT-GATED

Public-review automation with a safe submission boundary

App Store publish-back stays agent-approved; Google Play and Trustpilot require per-reply approval under canary. One public dispatch boundary, no automatic retry.

AUDIT COMPLETE

Support dashboard and metric-governance audit

Definitions and reporting logic reviewed across 2,262 saved questions. Live rollout remains operations/compliance approval-gated.

CONTROLLED LEARNING

Human-gated feedback and lesson promotion

Negative verdicts become candidates; a human approves what becomes a reusable, retrievable lesson.

Operational patterns

The queue is rarely the whole problem.

Support pressure is usually a systems problem in disguise: retrieval is weak, incidents are disconnected, manual work crosses too many tools, or an AI pilot has no reliable operating boundary.

01

The same questions are answered again and again.

Knowledge exists, but not in a form an agent can retrieve and trust while handling a live case.

02

Recurring incidents are investigated from zero.

History, observability, and ownership are scattered across systems, so diagnosis restarts every time.

03

Manual workflows cross too many systems.

The work is repetitive, but the failure boundary is unclear enough that naive automation makes it worse.

04

The AI pilot works—until it meets production.

Output quality is tested while permissions, retries, escalation, and human authority are left undefined.

Method

Diagnose first. Prototype cheaply. Promote what earns it.

Tools are chosen after the operational boundary is understood. The aim is evidence early, explicit control where failure is expensive, and a system a team can own.

Automate aggressively where failure is cheap. Constrain automation where failure is expensive.

Diagnose the operating system

Read the queue, incidents, data, ownership, and actual failure cost.

Prototype the smallest useful slice

Use the quickest honest path to test the operational hypothesis.

Prove behaviour, not just output

Evaluate retrieval, permissions, handoffs, failure states, and side effects.

Promote what proves itself

Move durable, high-value parts into tested and observable services.

Instrument and hand over

Leave status, stop conditions, and ownership legible after launch.

Where I add value

Useful when the pressure is repeatable.

The best work starts with a bounded operational failure: something real enough to inspect, but specific enough to improve without a grand transformation programme.

01

Queue diagnosis & automation design

Find repeat volume, handoff friction, and unsafe manual work before deciding whether deflection, routing, drafting, or better observability is the right answer.

02

Support knowledge & agent-facing AI

Build retrieval and assistance around real support artefacts, grounded answers, access boundaries, and the agent’s ability to correct the system.

03

Workflow reliability & observability

Harden an existing automation around idempotency, visibility, escalation, rollout gates, and a clear human decision boundary.

Public work

Work you can inspect directly.

These artefacts complement the confidential systems above. They show the reporting, source, and engineering discipline behind how I work; they are not substitutes for private proof.

01 / RESEARCH SYSTEM

Eye on Vietnam

A public regulatory tracker with sources, visible status, and a correction record. The point is not hot takes; it is traceable claims and honest uncertainty.

Read the research ↗
02 / PUBLIC ENGINEERING

Systems, tools, and working notes

Selected repositories and technical artefacts that represent the current quality bar—not a volume dump of every experiment.

View GitHub ↗

Background

Engineering shaped by support reality.

I learned systems from the point where they fail in front of users: support queues, provider integrations, logs, databases, incidents, handoffs, and incomplete operational context.

That perspective now carries through SQL and diagnostics into workflows, services, retrieval, and evaluation. The useful question is rarely “can an AI do this?” It is “what authority can it hold here, how do we know it behaved correctly, and who owns the exception?”

START WITHQueue data, incident history, failure cost, and ownership.
THEN BUILDBounded prototypes with explicit acceptance criteria.
CONTROLHuman decisions and deterministic boundaries around expensive mistakes.
LEAVE BEHINDObservable systems a real team can run and change.

Résumé

A human-readable résumé and a machine-readable version.

One page for a person; a plain-text version for an applicant tracking system. Both should say the same thing as the work above.

Contact

Tell me what your queue looks like.

For a role, a collaboration, or a bounded systems problem: a short note with useful context is more valuable than a vague “let’s connect.”

contact@oronculzac.comEmail the short brief ↗

Three useful things to include

  1. What keeps repeating or taking too long?
  2. Which systems, data, or team handoffs are involved?
  3. What would a better outcome change for the team?