The science behind safer driving
Six research components, each addressing a different cause of road accidents on long-distance bus routes in Sri Lanka, connected through a shared knowledge graph.
DriverGraphRAG — Multilingual Traffic-Law Assistant
A question-answering assistant that lets drivers, fleet managers and officers ask about traffic law and driver records in Sinhala, Singlish or English — and only answers with claims it can trace back to a knowledge graph.
The problem
General-purpose LLMs confidently recite traffic law from memory, can't bind Sinhala names to English graph records and give everyone the same answer regardless of who is asking.
Key capabilities
- Text and voice questions in Sinhala, Singlish and English (speech-to-text and text-to-speech)
- Neo4j knowledge graph linking Drivers, Trips, Routes, Violations and Motor Traffic Act sections
- LLM text-to-Cypher with cross-lingual schema & value linking (XLINK)
- Bridge retrieval that connects operational violations to the legal sections that govern them
- Layered hallucination control: an evidence-sufficiency gate, GAFL, Chain-of-Verification and conformal abstention
- Identity-aware personalisation and per-claim confidence rendered into the answer
What it produces
Vision-Based Fatigue Monitoring
Real-time drowsiness detection from a cabin camera, running face-landmark analysis directly in the browser so video never needs to leave the vehicle.
The problem
Long overnight routes make driver fatigue one of the leading causes of bus accidents, and it builds slowly enough that drivers rarely notice it themselves.
Key capabilities
- In-browser face landmark tracking with MediaPipe
- Eye Aspect Ratio (EAR) and blink-duration tracking for micro-sleep detection
- Mouth Aspect Ratio (MAR) for yawn detection
- Head-pose estimation to catch nodding and drooping
- Temporal buffering to smooth noisy frame-level signals
- Fusion with physiological and vehicle (CAN) signals for a combined fatigue score
What it produces
Driver Distraction Analysis
A deep-learning classifier that recognises distracting activities from cabin imagery and pairs them with the road scene to judge how risky the distraction is at that moment.
The problem
A glance at a phone on an empty straight road is very different from the same glance on a sharp bend — risk depends on both the driver and the road.
Key capabilities
- CNN classifier covering 10 cabin behaviours — safe driving, texting, phone calls, drinking, adjusting the radio, reaching behind and more
- Gaze estimation to judge whether eyes are on the road
- Road-scene awareness, including curve detection from the forward camera
- Risk engine that combines cabin state and road context
What it produces
RouteGuard — Driving Behaviour & Passenger Comfort
IMU-based telemetry that classifies driving style in real time, estimates passenger comfort and tracks component wear for predictive maintenance.
The problem
Harsh braking, aggressive cornering and sudden acceleration wear out vehicles and passengers alike, yet they are rarely measured on luxury long-distance buses.
Key capabilities
- Live accelerometer and gyroscope stream from an MPU motion sensor (with a simulator for demos)
- Jerk and feature engineering over 30-sample sliding windows
- Random Forest models for driving behaviour and passenger comfort
- Aggression meter, comfort gauge and real-time alert feed
- Route-zone breakdown and score comparison across trips
- Component-wear tracking against a fleet baseline for predictive maintenance
What it produces
Traffic-Violation Detection & Context-Aware Speed
Dash-camera and GPS analysis that detects traffic-light and pedestrian-crossing violations and judges speed against the real context — school zones, rain and route geometry.
The problem
A static speed limit ignores context. 50 km/h may be legal on paper but dangerous outside a school or in heavy rain.
Key capabilities
- YOLO-based detection of traffic lights and pedestrian crossings from dash-cam video
- GPS stop-compliance monitoring for red-light running and rolling stops
- Context-aware speed rules using live weather, school zones and the route corridor
- Voice alerts for the driver as violations happen
- Violations streamed into the knowledge graph through a real-time ingestion pipeline
What it produces
Driver Risk Scoring
A fair driver score that replaces flat point deductions with an exposure-normalised, Bayesian-smoothed and time-decayed risk rate.
The problem
Fixed deductions treat a driver with 5 trips the same as one with 5,000, never forgive old violations and ignore how severe each one was.
Key capabilities
- Risk expressed as a rate: severity per kilometre driven
- Bayesian prior pulls drivers with little history toward the fleet average
- Exponential time decay so violations age out and clean streaks raise the score
- Server-side severity grading instead of client-supplied deductions
- Explainable score breakdown, backtesting and half-life calibration
What it produces
A pipeline that earns trust
Black-box LLMs can't be fine-tuned to stop hallucinating, so the pipeline wraps the model with diagnosis before generation and verification after it.
- STEP 1AskVoice or text in Sinhala, Singlish or English. Speech is transcribed first.
- STEP 2GuardPrompt-injection, PII, toxicity, relevance and Cypher-injection guardrails screen the input.
- STEP 3UnderstandA LoRA-tuned Qwen2.5 model normalises Singlish, and XLINK binds names to graph values.
- STEP 4RetrieveText-to-Cypher and bridge retrieval pull operational and legal evidence from Neo4j.
- STEP 5DiagnoseThe evidence-sufficiency gate abstains before generation if the rows can't answer the question.
- STEP 6GenerateGemini writes a grounded answer, personalised with the caller's own graph profile.
- STEP 7VerifyGAFL, Chain-of-Verification and conformal abstention check every claim before it is shown.
Layered hallucination control
Each layer answers a different question about an answer, so their blind spots don't overlap. The gate is the only one that prevents a hallucination; the rest detect and correct it, and conformal calibration bounds how often the whole stack is wrong.
Built for Sri Lanka
Cross-lingual value linking lets a question like “කමල්ගේ දඩ මොනවද?” bind to the English record for Kamal in the graph — something plain text-to-Cypher can't do.
- 1
Evidence Sufficiency Gate
Asks whether the retrieved rows actually contain what this question needs, and abstains instead of letting the model fill gaps from memory.
Before generationPrevents - 2
GAFL — Graph-Anchored Faithfulness Ledger
Traces every number, name and section in the answer back to a retrieved row, and regenerates when values can't be found.
After generationDetects values - 3
CoVe — Chain-of-Verification
Turns the draft into verification questions to catch claims whose values are right but whose relationships are wrong.
After generationDetects relationships - 4
PGC — Parametric Grounding Contrast
Checks whether the model used the evidence or would have said the same thing from pre-trained memory.
After generationDetects memorisation - 5
Conformal Abstention
Calibrates a threshold that gives a statistical guarantee on how often an accepted answer is wrong.
Over the systemBounds the error
Publication · 2026
Diagnose, Verify, Personalize: A Black-Box GraphRAG Pipeline for Multilingual Driver-Safety and Traffic-Law Question Answering
Manuscript submitted for publication
