Skip to content
CITOS
Our research

The science behind safer driving

Six research components, each addressing a different cause of road accidents on long-distance bus routes in Sri Lanka, connected through a shared knowledge graph.

Component 01

DriverGraphRAG — Multilingual Traffic-Law Assistant

A question-answering assistant that lets drivers, fleet managers and officers ask about traffic law and driver records in Sinhala, Singlish or English — and only answers with claims it can trace back to a knowledge graph.

The problem

General-purpose LLMs confidently recite traffic law from memory, can't bind Sinhala names to English graph records and give everyone the same answer regardless of who is asking.

Neo4jGemini on Vertex AIBGE embeddingsQwen2.5-3B + LoRAFastAPIConformal prediction

Key capabilities

  • Text and voice questions in Sinhala, Singlish and English (speech-to-text and text-to-speech)
  • Neo4j knowledge graph linking Drivers, Trips, Routes, Violations and Motor Traffic Act sections
  • LLM text-to-Cypher with cross-lingual schema & value linking (XLINK)
  • Bridge retrieval that connects operational violations to the legal sections that govern them
  • Layered hallucination control: an evidence-sufficiency gate, GAFL, Chain-of-Verification and conformal abstention
  • Identity-aware personalisation and per-claim confidence rendered into the answer

What it produces

Grounded answer with cited legal sectionsPer-claim confidenceCertified abstention when evidence is insufficient
Component 02

Vision-Based Fatigue Monitoring

Real-time drowsiness detection from a cabin camera, running face-landmark analysis directly in the browser so video never needs to leave the vehicle.

The problem

Long overnight routes make driver fatigue one of the leading causes of bus accidents, and it builds slowly enough that drivers rarely notice it themselves.

MediaPipe Face LandmarkerWebSocketsTemporal fusionReact

Key capabilities

  • In-browser face landmark tracking with MediaPipe
  • Eye Aspect Ratio (EAR) and blink-duration tracking for micro-sleep detection
  • Mouth Aspect Ratio (MAR) for yawn detection
  • Head-pose estimation to catch nodding and drooping
  • Temporal buffering to smooth noisy frame-level signals
  • Fusion with physiological and vehicle (CAN) signals for a combined fatigue score

What it produces

Live fatigue levelBlink and yawn eventsFatigue alerts
Component 03

Driver Distraction Analysis

A deep-learning classifier that recognises distracting activities from cabin imagery and pairs them with the road scene to judge how risky the distraction is at that moment.

The problem

A glance at a phone on an empty straight road is very different from the same glance on a sharp bend — risk depends on both the driver and the road.

PyTorch CNNTransfer learningComputer visionFastAPI

Key capabilities

  • CNN classifier covering 10 cabin behaviours — safe driving, texting, phone calls, drinking, adjusting the radio, reaching behind and more
  • Gaze estimation to judge whether eyes are on the road
  • Road-scene awareness, including curve detection from the forward camera
  • Risk engine that combines cabin state and road context

What it produces

Distraction class & confidenceContext-aware risk level
Component 04

RouteGuard — Driving Behaviour & Passenger Comfort

IMU-based telemetry that classifies driving style in real time, estimates passenger comfort and tracks component wear for predictive maintenance.

The problem

Harsh braking, aggressive cornering and sudden acceleration wear out vehicles and passengers alike, yet they are rarely measured on luxury long-distance buses.

MPU IMU sensorRandom ForestSignal processingFastAPI streaming

Key capabilities

  • Live accelerometer and gyroscope stream from an MPU motion sensor (with a simulator for demos)
  • Jerk and feature engineering over 30-sample sliding windows
  • Random Forest models for driving behaviour and passenger comfort
  • Aggression meter, comfort gauge and real-time alert feed
  • Route-zone breakdown and score comparison across trips
  • Component-wear tracking against a fleet baseline for predictive maintenance

What it produces

Behaviour classComfort scoreMaintenance indicators
Component 05

Traffic-Violation Detection & Context-Aware Speed

Dash-camera and GPS analysis that detects traffic-light and pedestrian-crossing violations and judges speed against the real context — school zones, rain and route geometry.

The problem

A static speed limit ignores context. 50 km/h may be legal on paper but dangerous outside a school or in heavy rain.

YOLOv8UltralyticsGPS telemetryWeather APIGoogle Maps

Key capabilities

  • YOLO-based detection of traffic lights and pedestrian crossings from dash-cam video
  • GPS stop-compliance monitoring for red-light running and rolling stops
  • Context-aware speed rules using live weather, school zones and the route corridor
  • Voice alerts for the driver as violations happen
  • Violations streamed into the knowledge graph through a real-time ingestion pipeline

What it produces

Detected violationsLegal section mappingDriver alerts
Component 06

Driver Risk Scoring

A fair driver score that replaces flat point deductions with an exposure-normalised, Bayesian-smoothed and time-decayed risk rate.

The problem

Fixed deductions treat a driver with 5 trips the same as one with 5,000, never forgive old violations and ignore how severe each one was.

Bonus-MalusPoisson-Gamma modelBayesian smoothingFirestore

Key capabilities

  • Risk expressed as a rate: severity per kilometre driven
  • Bayesian prior pulls drivers with little history toward the fleet average
  • Exponential time decay so violations age out and clean streaks raise the score
  • Server-side severity grading instead of client-supplied deductions
  • Explainable score breakdown, backtesting and half-life calibration

What it produces

0–100 driver scoreRisk levelScore explanation
Deep dive · DriverGraphRAG

A pipeline that earns trust

Black-box LLMs can't be fine-tuned to stop hallucinating, so the pipeline wraps the model with diagnosis before generation and verification after it.

  1. STEP 1AskVoice or text in Sinhala, Singlish or English. Speech is transcribed first.
  2. STEP 2GuardPrompt-injection, PII, toxicity, relevance and Cypher-injection guardrails screen the input.
  3. STEP 3UnderstandA LoRA-tuned Qwen2.5 model normalises Singlish, and XLINK binds names to graph values.
  4. STEP 4RetrieveText-to-Cypher and bridge retrieval pull operational and legal evidence from Neo4j.
  5. STEP 5DiagnoseThe evidence-sufficiency gate abstains before generation if the rows can't answer the question.
  6. STEP 6GenerateGemini writes a grounded answer, personalised with the caller's own graph profile.
  7. STEP 7VerifyGAFL, Chain-of-Verification and conformal abstention check every claim before it is shown.

Layered hallucination control

Each layer answers a different question about an answer, so their blind spots don't overlap. The gate is the only one that prevents a hallucination; the rest detect and correct it, and conformal calibration bounds how often the whole stack is wrong.

Built for Sri Lanka

Cross-lingual value linking lets a question like “කමල්ගේ දඩ මොනවද?” bind to the English record for Kamal in the graph — something plain text-to-Cypher can't do.

  • 1

    Evidence Sufficiency Gate

    Asks whether the retrieved rows actually contain what this question needs, and abstains instead of letting the model fill gaps from memory.

    Before generationPrevents
  • 2

    GAFL — Graph-Anchored Faithfulness Ledger

    Traces every number, name and section in the answer back to a retrieved row, and regenerates when values can't be found.

    After generationDetects values
  • 3

    CoVe — Chain-of-Verification

    Turns the draft into verification questions to catch claims whose values are right but whose relationships are wrong.

    After generationDetects relationships
  • 4

    PGC — Parametric Grounding Contrast

    Checks whether the model used the evidence or would have said the same thing from pre-trained memory.

    After generationDetects memorisation
  • 5

    Conformal Abstention

    Calibrates a threshold that gives a statistical guarantee on how often an accepted answer is wrong.

    Over the systemBounds the error

Publication · 2026

Diagnose, Verify, Personalize: A Black-Box GraphRAG Pipeline for Multilingual Driver-Safety and Traffic-Law Question Answering

Manuscript submitted for publication

Request a copy