The problem
The project began with a list of 992 legislative records. Each record contained enough information to identify what needed to be reviewed, but not enough to compare which powers and functions appeared in the legislation, which institutions or officials were named, which sphere of government was implicated, how functions overlapped, or where the supporting text could be found. A list of legislation could not show where government authority sat or how functions crossed institutional and constitutional boundaries.
Context
The project supported a South African public-sector research team working on the allocation of powers and functions across government. The work was not legal advice and did not replace legal interpretation. It prepared structured research data for specialist review so the team could inspect, compare and refine the evidence with the source text close at hand.
Before and after
- Before: The starting point was a list of 992 legislative references with names, years, Act numbers and mixed law types across several decades. The corpus included principal legislation, amendment legislation and regulations, but there was no single source-linked database showing the powers, duties, institutions, officials, spheres or supporting text inside those documents.
- After: The material moved through one controlled route: legislative references to official source retrieval, PDF and HTML parsing, stable law and source IDs, draft power-and-function extraction, law-level theme assessment, evidence linking, QA review and reusable research outputs.
Constraints
A legislative record is not one analytical record. One Act or regulation may contain several separate powers, different responsible institutions, different responsible officials, duties and discretionary powers, shared or split functions, national, provincial and municipal dimensions, scheduled and residual elements, fiscal provisions, assignments, delegations, authorisations and issues requiring specialist or case-law review. The workflow needed to separate factual source fields from draft interpretive mappings and review fields.
What the team needed
- A law and source register that connected each supplied legislative record to the official source material used for review
- A schema that could separate document metadata, draft powers-and-functions mappings, constitutional classifications, review flags and evidence locators
- A way to break long legislative documents into discrete analytical records without treating each document as one finding
- A review layer for uncertain mappings, possible hidden assignments, split functions and records requiring specialist legal or policy judgement
- A bounded retrieval layer that could help researchers query the project material without treating AI output as legal advice
- A reusable data layer for research packs, reports, presentations and web-based evidence views
If a research team needs to prove where each finding came from, test the source-traceability risk before the database becomes a report or public-facing output.
What I built
A source-linked legislative evidence workflow that moved from a supplied research corpus to official source documents, parsed text, draft power-and-function mapping records, evidence extracts, law-level theme assessments, QA review queues, a structured research data pack, a bounded-source AI research assistant and an interactive evidence atlas.
Named systems and workflow pieces
- A law register covering 992 legislative records in the supplied research corpus
- A source register linked to official legislative source documents
- PDF and HTML parsing workflows for preparing source text for structured extraction
- Stable law IDs, source IDs and evidence IDs across the workflow
- A powers-and-functions schema covering identification, administration, responsible institutions, officials, spheres, Schedule 4 or 5 classification, residual classifications, power type, function role, assignments, delegations, authorisations, fiscal flags and review fields
- 3,385 draft power-and-function mapping records, with a new row created whenever a distinct power, duty, responsibility or role appeared
- 2,976 law-level assessments across three research questions
- 3,765 source-linked evidence extracts supporting mappings and assessments
- A QA review queue for records needing legal, policy or methodological judgement
- A structured research data pack with Excel and CSV exports, field notes and usage notes
- A bounded-source AI research assistant configured around the project databases and legislative source pack
- An interactive evidence atlas with linked views for legislation, powers and functions, schedules, departments, sectors, analytical themes, sources and downloads
Where this connects to the services
This case study sits mainly under Traceable Evidence Workflow Support because the project needed source registers, parsed documents, evidence extracts, draft mappings, review flags, QA queues and source traceability. It also connects to Data Use, Reporting & Communication Systems because the same reviewed evidence layer feeds a structured research data pack, bounded-source research assistant, interactive evidence atlas, reports and presentation material.
One route from legislative references to reviewable research outputs
The workflow kept each draft mapping close to the source text while turning a long legislative corpus into structured data, review queues and reusable outputs.
Each supplied legislative record was linked to official source material and stable law and source IDs.
PDF and HTML material was prepared so long documents could be broken into structured analytical records.
Draft powers, duties, institutions, officials, spheres and classifications were captured as discrete mapping records.
Evidence IDs linked mappings and theme assessments back to supporting legislative extracts.
Uncertain or interpretive records were routed into a QA queue for specialist human judgement.
The same evidence layer fed the data pack, bounded-source research assistant and interactive evidence atlas.
Legislative record to source document to evidence extract to draft mapping to review status to research pack, assistant and atlas
How it worked
The workflow moved from raw material to usable output through a short sequence of controlled steps.
Process
- 01
Created a law register from the supplied corpus so each legislative record had a stable identifier and basic metadata.
- 02
Searched for official source material corresponding to each supplied record and created a source register linked to the law register.
- 03
Stored and parsed PDF or HTML source material so the legislative text could be prepared for structured extraction.
- 04
Designed the powers-and-functions schema before extraction began, separating source fields, draft mapping fields, research fields and review fields.
- 05
Broke legislative text into discrete draft mapping records for powers, duties, responsibilities and roles named or implicated by the source material.
- 06
Assessed each legislative record against three research questions: scheduled or residual, provincial or municipal dimension, and possible hidden assignment.
- 07
Linked analytical records back to evidence extracts so researchers could move from a filtered database record to the supporting legislative text.
- 08
Routed uncertain or interpretive records into a QA queue with confidence, review status, reviewer notes and specialist-review flags.
- 09
Prepared a structured research data pack, bounded-source AI assistant and interactive evidence atlas from the same evidence layer.
Outputs
These were the named assets, dated deliverables, and working materials left behind by the project.
Working outputs
- Law register with 992 legislative records
- Source register and source document corpus
- Evidence extract table with 3,765 source-linked excerpts
- Draft powers-and-functions mapping table with 3,385 records
- Theme assessment table with 2,976 law-level assessments
- QA review queue for uncertain and interpretive mappings
- Processing log and workflow notes
- Structured Excel workbook and separate CSV exports
- Field and usage notes for the research team
- Bounded-source AI research assistant
- Interactive evidence atlas and research portal interface
Current project status
The first data-build phase has produced a source corpus, eight linked database and workflow tables, 3,385 draft power-and-function mappings, 2,976 law-level theme assessments and 3,765 evidence extracts. The next phase focuses on specialist review, data refinement, further analysis and report development.
What has been completed so far
- Created one searchable evidence base instead of repeated document-by-document searching
- Gave the research team a granular view of powers and functions within each legislative record
- Made it easier to compare responsibilities across spheres, institutions, officials and department or sector categories
- Improved visibility of split functions, overlapping responsibilities and possible hidden assignments without treating them as final legal conclusions
- Created a structured route for specialist review through confidence fields, human-review flags, QA queues and reviewer notes
- Improved retrieval of supporting legislative text for reports, presentations and further analysis
- Kept a clear distinction between first-pass extraction, draft research mappings and approved findings
Legislative document extraction only becomes useful when the database design, evidence links, review fields and output layer make the material checkable. The goal is not to turn draft mappings into legal conclusions, but to give the research team a structured evidence base for specialist review.
What this proves
- Sourcing and organising a large body of official public-sector source material
- Designing a granular database before extraction begins
- Breaking long legislative documents into discrete analytical records
- Mapping institutions, officials, spheres, functions and constitutional classifications as draft research data
- Preserving the route from each mapping back to supporting legislative text
- Routing uncertain records into human review instead of hiding uncertainty
- Turning one structured evidence layer into a data pack, research assistant and interactive evidence atlas
- Working as a specialist subcontractor inside a larger public-sector research team
Best fit
These are the situations where this kind of evidence workflow tends to be the strongest fit.
Who this is best for
- Public-sector research projects working with large legislative, policy or regulatory source bases
- Policy and legislative reviews that need source-linked evidence behind every finding
- Powers-and-functions studies across spheres, institutions, officials and government sectors
- Research teams that need to turn official documents into structured reviewable data
- Projects where AI-assisted retrieval must remain bounded by approved sources and human review
- Teams building research packs, dashboards, portals or evidence atlases from the same structured data layer
Service stack connected to this case study
This case study sits inside the same delivery work, service logic, and practical outcomes shown across the site.
Turn interviews, submissions, case studies, survey comments, documents, and field notes into coded evidence, quote banks, synthesis tables, findings, recommendations, and report-ready outputs.
Use structured data in reports, dashboards, internal tools, public microsites, applications, presentations, annual reports, and decision-support workflows.
Traceable Evidence Workflow Support
This service packages the same kind of source register, document parsing, evidence database, review queue and source-linked output workflow for evidence-heavy research teams.
A route for legislative and policy evidence work
Use this route when legislation, policy documents, public-sector source material or research records need to become structured, source-linked evidence that can support review and reporting.
View Traceable Evidence Workflow Support