Extraction rules, taxonomies, labelling conventions, quality contracts, permissions and exception handling are structured in a single operational foundation.
Apshan
For Apshan, we designed a lakehouse that turns heterogeneous sources into structured, quality-controlled data available to agents. Custom CLIs run the routine flows autonomously; anomalies and rule changes remain subject to human approval.
The system delivered
Operations handled
Architecture delivered
- 01Custom CLIExtraction and normalisation
Specialised commands for each source and each stage of the pipeline.
- 02AWS S3Canonical object storage
- 03Apache IcebergVersioned, interoperable tables
- 04LakekeeperApache Iceberg catalogue
- 05QdrantVector memory for agents
- 06DagsterOrchestration and observability
- 07AWS IAMIdentities and least-privilege access
- 08A2A ProtocolInteroperability between agents
Connected tools
- AWS S3
- AWS IAM
Operation and control
Operational memory
Structured operational knowledge
Reusable corrections
A validated human correction becomes a normalisation rule, a labelling example or a control reusable in later runs. The system never changes its rules on its own.
Operational safeguards
- Data qualityPublication blocked if controls fail
- Every batch must meet the schema and quality contracts before entering the tables available to agents.
- Access controlLeast-privilege rights per service
- AWS IAM separates the identities and limits each component to the resources it needs.
- ObservabilityObservable runs
- The orchestration exposes pipeline state, failures and the steps that need intervention.
- RecoveryVersioned data and tables
- Canonical storage and Apache Iceberg snapshots make it possible to recover an earlier state of the data.
- AuditabilityTraceable pipeline decisions
- Validations, rejections and exception handling stay linked to the run that produced them.
Case study
The challenge
Apshan needed to collect and reconcile fashion data from multiple sources, with different formats, taxonomies and levels of quality. Manual processing could neither keep up with the volume nor produce a coherent memory for the products and the agents.
The system had to automate extraction and labelling without letting non-conforming data through, while keeping explicit access rights and human validation points for exceptions.
The solution
We delivered a unified data foundation: canonical storage on AWS S3, Apache Iceberg tables, a Lakekeeper catalogue and Qdrant vector memory. Dagster orchestrates the flows and makes every run observable.
- Custom CLIs extract, normalise, label and enrich the data.
- Quality gates block non-conforming batches from publication and surface the exceptions.
- AWS IAM limits access per service; the A2A protocol makes the memory and capabilities available to authorised agents.
- Apshan's bilingual website completes the system as the platform's public showcase.
Public interface, also delivered
Which part of your client work would you hand over first?
Book a call to map how that work moves through your agency today, agree the first workflow to build, and set what stays under your control.







