Technical portfolio, revised 6 Oct 2026 Hà Nội, GMT+7 Taking freelance projects

Keras Vũ

AI engineer. I build retrieval, vision and speech systems, and I keep them running after the demo.

Right now: RAG chatbots with citations

Table 1. Selected results from shipped work.
0/50legal questions answered from the right source document
~0 msfrom camera frame to alert on an edge AI box
0%traffic-sign Q&A after LoRA fine-tuning Qwen2.5-VL-7B
0 GBof GPU memory for a full ASR → translation → TTS dub

Recorded from running systems

Two recordings from the staging environment of the multi-tenant AI platform where I am the only AI engineer. Nothing here is a mockup.

REC AI Box 3 · Camera 123 00.0 s

HIGH New AI event

Intrusion · restricted zone

aibox-box3 · Camera 123 · yolo11m.engine

conf 0.81 · frame → event 71 ms

Fig. 1. An intrusion alert on AI Box 3, recorded 5 Oct 2026. The box and scores are the ones stored with the event, and the alert appears at the moment the event was raised. Faces are pixelated.

A person walks into a restricted zone

The AI box at the site reads the camera's RTSP stream and runs a YOLO11 TensorRT engine. When someone crosses the zone drawn for this camera, the box raises an event and the web portal receives it with a snapshot and the detection box.

Frame to event
71 ms
Frame to server
39 ms
Rules fired
intrusion 0.81, no PPE 0.89

Event feed

  • Waiting for the first event…
Document assistant 0.0 s
Fig. 2. A replay of a real answer from 6 Oct 2026, with the exact text and the timing measured in the browser.

A question answered from a technical report

I uploaded the Qwen3 Technical Report as a PDF and asked how the model mixes thinking and non-thinking modes. The answer took 9.5 seconds and cites the sections it used.

  1. Analyse the question.
  2. Hybrid search: bge-m3 dense vectors and BM25, fused with RRF in Qdrant.
  3. Rerank the passages with a cross-encoder.
  4. Answer with self-hosted Qwen on vLLM, streamed with citations.
Screenshot of the same answer in the portal
The same answer in the portal.
Document library with indexing status
Document library with indexing status.

Systems I have built

From retrieval and fine-tuning to edge deployment and speech. What each one does, and what it runs on.

Document Assistant

RAG · LLM · SaaS

A chat assistant inside a multi-tenant SaaS platform. Teams upload their own files and ask questions in English or Vietnamese, and every answer cites the section it came from.

  • Upload PDF, Word, Markdown or TXT. Indexing runs in the background.
  • Structure-aware chunking keeps headings and sections together.
  • Hybrid search with bge-m3 dense vectors and BM25 in Qdrant.
  • Streaming answers with numbered citations linked to the passage.
  • Each tenant sees only its own documents.
  • Self-hosted Qwen on vLLM, so no document goes to a third-party API.
  • Runs on a central GPU server or on a box at the customer's site.

Qwen · vLLM · Qdrant · bge-m3 · BM25 · FastAPI · Docker · K3s

Assistant answering about the Qwen3 report with citations
Fig. 3. English answer with numbered citations.

Vietnamese Legal Q&A

RAG · AI agent · Legal

Question answering over Vietnamese laws and decrees. Answers point to the exact article and clause, not just a page.

  • Parses laws into their Chapter → Article → Clause → Point tree.
  • 22,000+ chunks, each carrying its full legal path for citations.
  • A LangGraph agent rewrites the question and splits it into sub-queries.
  • Hybrid dense and BM25 search with RRF fusion.
  • Cross-checks related documents, such as a decree against its law.
  • On a 50-question test set: right document 44 times, exact clause 40 times.

LangGraph · Qwen · vLLM · Qdrant · bge-m3 · BM25 · FastAPI

Legal answer about personal data breach notification
Fig. 4. The 72-hour breach notice in Decree 356/2025.
Citations to articles and clauses
Fig. 5. Citations by article and clause.

Edge AI Video Analytics

Computer vision · Edge AI · TensorRT

Cameras at each site stream to an AI box on the premises. Detectors run there, and events reach a central portal with a snapshot about 71 ms after the frame.

  • Multi-tenant portal: sites, AI boxes, cameras and users per customer.
  • Intrusion detection with restricted zones drawn per camera.
  • PPE detection for workers missing protective gear.
  • Every event stored with snapshot, clip, boxes, severity and latency.
  • Dashboard with events per day and filters by site, camera and severity.
  • Licensed models packaged as TensorRT engines and pulled to the boxes.

TensorRT · YOLO · RTSP · K3s · FastAPI · RabbitMQ · Docker

Event snapshot with restricted zone and flagged person
Snapshot, face blurred.
Intrusion event detail page
Event detail with latency.
Dashboard overview
Dashboard.
Licensed models list
Licensed TensorRT models.

People & Vehicle Access

Face recognition · ANPR · VLM

Access control at a gate, on one edge node per site. I was the only AI engineer, from system design to deployment.

  • Face registration and recognition from an enrollment camera.
  • License plate detection and recognition.
  • A review service where operators confirm or reject uncertain matches.
  • A Qwen vision-language model on vLLM double-checks results.
  • Single-node K3s for four RTSP feeds and the enrollment camera.

Face recognition · LPR/OCR · Qwen VLM · vLLM · RabbitMQ · K3s

4 RTSP cameras Enrollment cam Edge node · K3s · GPU Face recognition License plate OCR Qwen VLM check (vLLM) RabbitMQ workflows Backend& review UI
Fig. 6. Services on one edge node.

Video Translation & Dubbing

Speech · TTS · LLM

A web app I built on my own. Upload an English or Chinese video and get Vietnamese subtitles and a Vietnamese voice-over, all on one 8 GB GPU.

  • Speech recognition with Qwen ASR, OCR for burned-in subtitles.
  • Translation with Gemini that keeps sentence timing.
  • Local text-to-speech with VieNeu-TTS and OmniVoice.
  • Timing alignment so speech fits each subtitle slot, with automatic checks.
  • A 264-second video exports in about 14 minutes on an RTX 3060 Ti.

Qwen ASR · OCR · Gemini · VieNeu-TTS · OmniVoice · FFmpeg · ONNX Runtime

Video in ASR + OCR Translate TTS Timing QA FFmpeg Dubbed video
Fig. 7. The dubbing pipeline.

LLM / VLM fine-tuning

Traffic-sign Q&A with Qwen2.5-VL

Fine-tuned Qwen2.5-VL-7B with LoRA to answer multiple-choice questions about traffic-sign images.

0%1,437 of 1,500 test questions

Voice AI · Function calling · IoT

A voice assistant on an ESP32-S3

A small device you talk to. It sends speech to a server, an LLM decides what to do, and the device answers out loud. Function calling covers weather, news and music.

Vision models I have trained

Measured on independent QC test sets, not on the training split.

TaskResultNotes
No-helmet detection98.5% / 95.7%Precision and recall, running in real time inside a C++ application.
Drowsiness detection70% → 96%Classroom video. Most of the gain came from fixing data balance and class definitions, not a bigger model.
Violence detection93.2%Class precision with MoViNet and frame voting on classroom video.
Talking detection91%Eye tracking combined with object detection.
Face anti-spoofing80% → 87.5%Improved by generating augmented spoof examples.
Text-to-person search~1M peopleFine-tuned SigLIP2 to find people from descriptions like “man in a red jacket carrying a backpack”.
Table 2. Vision models and how they were measured.

Tools I use every week

LLM & RAG
Qwen, vLLM, LangGraph, LangChain, Qdrant, bge-m3, BM25, LoRA
Vision
PyTorch, YOLO, OpenCV, SigLIP, MoViNet, NVIDIA DeepStream
Speech
Qwen ASR, VieNeu-TTS, OmniVoice, FFmpeg
Speed
ONNX Runtime, TensorRT, FP16 / INT8, CUDA, NVIDIA Triton
Serving
Python, C++, FastAPI, RabbitMQ, Docker, Kubernetes (K3s), MinIO

Where I have worked

  1. Apr 2026 – now

    AI Engineer, Smartech

    The only AI engineer in a four-person product team: document assistant, edge video analytics platform, people and vehicle access management.

  2. Mar 2025 – Mar 2026

    Computer Vision Engineer, AIBOX

    Vision and multimodal R&D: classroom video analytics on DeepStream (drowsiness, talking and violence detection) and SigLIP2 fine-tuning for text-based person retrieval.

  3. 2021 – 2026

    Hanoi University of Science and Technology

    Bachelor's degree in Control Engineering and Automation.

Have documents, images, video or audio you want AI to work on?

I take projects through Upwork, from a first prototype to a system running on your own servers.

View my Upwork profile