- Professional work · AEM Energy Solutions Sdn Bhd
- Project number
- Project / 002
- Year
- 2025
- Category
- Data Engineering
- Status
- Proof of concept
QualityPlus
Data quality management, turned into a product. Score a dataset, fix what failed inside a governed workflow, and publish a version everyone can trust.
- Role
- Solo: product, engineering, AI workflows
- Domain
- Oil & gas datasets
- Stack
- React · TypeScript · Supabase · n8n · Ollama
- Repository
- Private
- Status
- Proof of concept · demonstrated at SPE AI Summit
01 /Overview
Measure it. Fix it. Trust it.
QualityPlus lets a team upload a dataset, run configurable checks across four quality dimensions, and work every failed cell through a review workflow until a corrected “trusted” version is published.
02 /The problem
Bad data gets fixed in spreadsheets.
Operational datasets arrive incomplete, duplicated or inconsistent. The fixes usually happen by hand in a copy of a spreadsheet, with no record of what changed, who approved it, or which rules were applied.
Every project also ends up reinventing its own validation rules, so the same mistakes get caught differently, or not at all, depending on who looks.
03 /The idea
Quality as a workflow, not a report.
- Assessment
- Remediation
- Approval
- Publish
A score on its own is not enough; people need a way to act on it. So the score becomes the start of a workflow. AI can suggest at each step, but a person always decides.
04 /System architecture
Scoring runs in the browser. AI stays optional.
Step 1: SOURCE
Dataset upload
Dataset files uploaded into a project.
Step 2: PROCESS
Quality engine
Pure TypeScript; scores without a database round-trip.
Step 3: DATABASE
Supabase
Scores history, templates, corrections, auth, roles.
Step 4: AGENT
n8n workflows
Three webhooks: rule check, summary, remediation.
Step 5: MODEL
Ollama
An open model served through Ollama, called from the n8n workflows.
Step 6: OUTPUT
Trusted data
Corrected dataset + PDF quality report.
05 /Data / intelligence
Four dimensions, three AI touchpoints.
Completeness
Selected columns must be present and non-empty.
Uniqueness
Single-column or composite-key uniqueness across rows.
Validity
Ranges, thresholds, allowed values, regex and data types, seeded by automatic type inference.
Consistency
Values must match a reference list, an uploaded file, or another dataset already in the platform.
01AI rule check
Reads column names against an O&G domain knowledge base, suggests validity rules, and flags columns that should be consistency checks instead.
02AI summary
Writes a plain-language summary of a saved score: key issues, links to the failing rows, recommendations.
03AI remediation
Suggests a replacement for each failed cell, with a reason and confidence, using passing values from the same column as style examples.
06 /Build
Keep the AI’s answer separate from the human’s.
01Traceable corrections
Every fix records whether it was manual, an accepted AI suggestion or a modified one, keeping both the AI’s value and the final value.
02Roles
Editors submit corrections; owners and co-owners approve or reject, with a required reason.
03Governed templates
A check configuration becomes a template that is reviewed and approved before other projects can reuse it.
04Derived state
Workflow phase is computed from failed, addressed and pending counts rather than stored, so it cannot drift.
05A pivot
An early Python backend on Render was replaced by talking to Supabase directly. It was simpler to run and there was less to break.
08 /Outcome
Live, and still evolving.
- Quality dimensions
- 4
- AI workflows
- 3
- all optional
- Workflow phases
- 4
- assess → publish
- Access roles
- 4
- owner → viewer
The remediation workflow is the most actively developed part: dataset refresh, visual marking of corrected cells and PDF export were the latest additions.
09 /What I learned
Suggest, do not decide.
01Separate suggestion from decision
Storing the AI’s value next to the human’s made the system auditable, and made the AI easier to evaluate.
02Keep the model swappable
The app only talks to n8n webhooks, so where the model runs is a configuration choice in n8n, not a change to the app.
03Governance is a feature
Rules only get reused once there is a review step people trust.
Read the journal
Where AI Belongs in a Data Quality Workflow10 /Next project
Duplicate Question Detection (NLP)Master's research: FastText embeddings plus PCA to detect duplicate questions on 404k Quora pairs, keeping 83.68% accuracy while cutting training time by two thirds and memory by 83%.

