df5e2b3e78
- MVP Milestone: 'First Verified Decision'
- Developer films with quiXzoom, imports to Lab, corrects AI,
creates Decision Case, follows chain with full traceability
- When this works = first complete verifiable Control Intelligence pipeline
- Three development phases:
Phase 1 (Essential): Ingestion, Dataset Explorer, Annotation, Decision Case Viewer
Phase 2 (Scale): Replay, Benchmark, Evaluation
Phase 3 (Advanced): GPU Jobs, Hyperparameter Runs, Model Promotion, Canary
- Product Architecture: quiXzoom → Observations → Intelligence Lab →
Improved Models → LandveX → Better Decisions → Feedback → Intelligence Lab
- Two products: quiXzoom (observations), LandveX (decisions)
- Intelligence Lab = the factory that improves both
- New areas:
- Data Quality: Healthy/Blurred/Duplicate/Wrong GPS/Night/Rain/Occluded
+ Coverage (Roads, Buildings, Signs, Drainage, Vegetation)
- Decision Analytics: Acceptance Rate, Ignore Rate, Accuracy,
Insufficient Evidence, Data Collection Value
- Business value metrics, not traditional AI metrics
Rationale: Build MVP first, prove first real workflow, then scale.
Decision Cases are the heart. Data Quality explains model performance.
Decision Analytics measure business value.
476 lines
11 KiB
Markdown
476 lines
11 KiB
Markdown
# LANDVEX INTELLIGENCE LAB
|
|
|
|
**Internal Development Environment for Control Intelligence**
|
|
|
|
| | |
|
|
|---|---|
|
|
| **Version** | 1.0 |
|
|
| **Status** | SPECIFICATION |
|
|
| **Purpose** | Build, test, and verify LandveX Control Intelligence before production |
|
|
|
|
---
|
|
|
|
## Core Principles
|
|
|
|
**LandveX Intelligence Lab does not produce AI models. It produces verified Control Intelligence.**
|
|
|
|
This is an internal tool. Never a customer product.
|
|
|
|
**All artifacts are immutable and versioned.**
|
|
- Raw images are never modified
|
|
- Annotations are versioned
|
|
- Models are versioned
|
|
- Evaluation reports are versioned
|
|
- Decision Cases are versioned
|
|
- Replay results are saved as new artifacts
|
|
|
|
**Every change must be traceable to a specific experiment, model version, dataset, and decision.**
|
|
|
|
## Product Architecture
|
|
|
|
```
|
|
quiXzoom
|
|
↓
|
|
Observations
|
|
↓
|
|
LandveX Intelligence Lab
|
|
↓
|
|
Improved Models
|
|
↓
|
|
LandveX
|
|
↓
|
|
Better Decisions
|
|
↓
|
|
Feedback
|
|
↓
|
|
Intelligence Lab
|
|
```
|
|
|
|
**Two products:**
|
|
- **quiXzoom** — produces observations
|
|
- **LandveX** — produces decisions
|
|
- **Intelligence Lab** — the factory that improves both
|
|
|
|
---
|
|
|
|
## Repository
|
|
|
|
```
|
|
landvex-intelligence-lab
|
|
```
|
|
|
|
Separate from:
|
|
- `quixzoom-app`
|
|
- `landvex-web`
|
|
- `aamos-core`
|
|
|
|
---
|
|
|
|
## Architecture
|
|
|
|
```
|
|
LandveX Intelligence Lab
|
|
│
|
|
┌─────────────────┼─────────────────┐
|
|
│ │ │
|
|
Developer UI Training API Experiment API
|
|
│ │ │
|
|
└─────────────────┴─────────────────┘
|
|
│
|
|
Event Bus / Queue
|
|
│
|
|
─────────────────────────────────────────────────────
|
|
Storage Layer
|
|
Images │ Videos │ Missions │ Models │ Logs
|
|
S3/R2 Buckets + PostgreSQL + Neo4j
|
|
─────────────────────────────────────────────────────
|
|
│
|
|
Processing Pipelines
|
|
```
|
|
|
|
## Navigation
|
|
|
|
```
|
|
Dashboard
|
|
├── Models
|
|
├── Datasets
|
|
├── Annotations
|
|
├── Training
|
|
├── Evaluation
|
|
├── Decision Cases
|
|
├── Replay
|
|
├── Validation
|
|
├── Promote Model
|
|
└── Settings
|
|
```
|
|
|
|
---
|
|
|
|
## Dashboard
|
|
|
|
Shows AI status, not business data.
|
|
|
|
```
|
|
┌─────────────────────────────────────┐
|
|
│ Datasets │
|
|
│ Healthy: 12 │
|
|
│ Needs Review: 3 │
|
|
│ Corrupted: 0 │
|
|
├─────────────────────────────────────┤
|
|
│ Training Jobs │
|
|
│ Running: 2 │
|
|
│ Queued: 1 │
|
|
│ Completed: 47 │
|
|
│ Failed: 0 │
|
|
├─────────────────────────────────────┤
|
|
│ Decision Cases │
|
|
│ Validated: 23 │
|
|
│ Pending: 5 │
|
|
│ Rejected: 2 │
|
|
├─────────────────────────────────────┤
|
|
│ Replay Jobs │
|
|
│ Ready: 8 │
|
|
│ Running: 1 │
|
|
│ Finished: 34 │
|
|
└─────────────────────────────────────┘
|
|
```
|
|
|
|
---
|
|
|
|
## Mission Replay
|
|
|
|
Click through the entire chain:
|
|
|
|
```
|
|
Video → Frame → Bounding boxes → Detected objects → Evidence → Finding → Decision → Business Impact
|
|
```
|
|
|
|
---
|
|
|
|
## Annotation
|
|
|
|
```
|
|
┌─────────┬─────────────┬──────────────┐
|
|
│ Video │ AI Suggestion│ Manual │
|
|
│ │ │ Correction │
|
|
├─────────┼─────────────┼──────────────┤
|
|
│ │ Object: │ Correct? │
|
|
│ │ Road Crack │ YES / NO │
|
|
│ │ Confidence: │ │
|
|
│ │ 82% │ Severity: │
|
|
│ │ │ Low / Medium │
|
|
│ │ │ / High │
|
|
└─────────┴─────────────┴──────────────┘
|
|
```
|
|
|
|
---
|
|
|
|
## Decision Cases
|
|
|
|
**The most important asset.**
|
|
|
|
Not images. Not videos. Not AI models.
|
|
|
|
But:
|
|
```
|
|
Observation → Evidence → Finding → Decision → Outcome → Learning
|
|
```
|
|
|
|
After a few years, hundreds of thousands of verified Decision Cases.
|
|
Not just a training dataset — a library of how real observations lead to real decisions and real outcomes.
|
|
|
|
```
|
|
Case #4232
|
|
├── Reality
|
|
├── Observation
|
|
├── Evidence
|
|
├── Finding
|
|
├── Decision
|
|
├── Outcome
|
|
└── Learning
|
|
```
|
|
|
|
All cases playable.
|
|
|
|
---
|
|
|
|
## Benchmark
|
|
|
|
Compare models:
|
|
|
|
| Model | Precision | Recall | F1 | Latency | Decision Accuracy |
|
|
|-------|-----------|--------|----|---------|-------------------|
|
|
| YOLO v8 | 0.89 | 0.87 | 0.88 | 45ms | — |
|
|
| Grounding DINO | 0.91 | 0.85 | 0.88 | 120ms | — |
|
|
| SAM | 0.88 | 0.90 | 0.89 | 200ms | — |
|
|
| Custom | 0.92 | 0.91 | 0.915 | 60ms | 0.87 |
|
|
|
|
---
|
|
|
|
## Replay
|
|
|
|
Find regressions:
|
|
|
|
```
|
|
Mission 213
|
|
├── Play
|
|
├── Show AI
|
|
├── Show Human Annotation
|
|
├── Differences
|
|
├── New Model
|
|
└── Old Model
|
|
```
|
|
|
|
---
|
|
|
|
## Validation
|
|
|
|
```
|
|
Field Trials
|
|
Scenario Tests
|
|
Decision Tests
|
|
Evidence Tests
|
|
Golden Failures
|
|
Regression Tests
|
|
```
|
|
|
|
---
|
|
|
|
## Promote Model (Not Deploy)
|
|
|
|
```
|
|
Development → Validation → Pilot → Production
|
|
```
|
|
|
|
Not "Deploy". "Promote Model".
|
|
|
|
---
|
|
|
|
## Experiments
|
|
|
|
```
|
|
Experiments
|
|
├── EP-1.0
|
|
├── DS-001
|
|
├── DS-002
|
|
├── DS-003
|
|
├── Field Trials
|
|
└── Metrics
|
|
```
|
|
|
|
Link experiment protocol to real development and validation data.
|
|
|
|
---
|
|
|
|
## Processing Pipelines
|
|
|
|
### 1. Ingestion Pipeline
|
|
|
|
**Input:** Images, video, GPS, EXIF, metadata
|
|
**Checks:** Checksums, versioning
|
|
**Output:** Raw Dataset
|
|
|
|
### 2. Dataset Pipeline
|
|
|
|
**Input:** Raw Dataset
|
|
**Checks:** Sort, deduplicate, quality control, resolution, blur detection, GPS validation
|
|
**Output:** Validated Dataset
|
|
|
|
### 3. Annotation Pipeline
|
|
|
|
**Input:** Validated Dataset
|
|
**Process:** AI suggestions, manual correction, label versions, consensus
|
|
**Output:** Verified Dataset
|
|
|
|
### 4. Training Pipeline
|
|
|
|
**Input:** Verified Dataset
|
|
**Process:** Start training, hyperparameters, checkpoints, GPU jobs
|
|
**Output:** Model Artifact
|
|
|
|
### 5. Evaluation Pipeline
|
|
|
|
**Input:** Model Artifact
|
|
**Metrics:** Precision, recall, F1, decision accuracy, regression
|
|
**Output:** Evaluation Report
|
|
|
|
### 6. Replay Pipeline
|
|
|
|
**Input:** Old missions, Model v14, Model v15
|
|
**Process:** Run both models, compare differences
|
|
**Output:** Regression Report
|
|
|
|
### 7. Decision Validation Pipeline
|
|
|
|
**Input:** Observation
|
|
**Process:** Full chain — Observation → Evidence → Finding → Decision → Business Impact
|
|
**Output:** Decision Validation Report
|
|
|
|
---
|
|
|
|
## Bucket Structure
|
|
|
|
```
|
|
raw-images/
|
|
raw-video/
|
|
missions/
|
|
datasets/
|
|
annotations/
|
|
training/
|
|
models/
|
|
replays/
|
|
evaluation/
|
|
decision-cases/
|
|
field-trials/
|
|
exports/
|
|
archives/
|
|
```
|
|
|
|
All content is versioned:
|
|
```
|
|
model-v14/
|
|
model-v15/
|
|
model-v16/
|
|
```
|
|
|
|
---
|
|
|
|
## AI Job Queue
|
|
|
|
All jobs are asynchronous:
|
|
|
|
```
|
|
Upload → Queue → Worker → GPU → Storage → Notification
|
|
```
|
|
|
|
Not synchronous API calls.
|
|
|
|
---
|
|
|
|
## What This Tool Collects
|
|
|
|
- Model training
|
|
- Annotation
|
|
- Datasets
|
|
- Replay
|
|
- Decision chains
|
|
- Validation
|
|
- Regression tests
|
|
- Experiments
|
|
- Model promotion
|
|
|
|
---
|
|
|
|
## New Developer Experience
|
|
|
|
A new AI engineer should open the repo and within minutes understand:
|
|
|
|
**"This is the tool where we build, test, and verify LandveX Control Intelligence before anything reaches production."**
|
|
|
|
---
|
|
|
|
## Relationship to Principles
|
|
|
|
- All development in Git
|
|
- All experiments reproducible
|
|
- All models traceable from training to validation to production
|
|
- Version control and traceability
|
|
|
|
---
|
|
|
|
## MVP Milestone: "First Verified Decision"
|
|
|
|
**Definition:**
|
|
A developer can:
|
|
1. Film a real object with quiXzoom
|
|
2. Import material to Intelligence Lab
|
|
3. Review and correct AI interpretation
|
|
4. Create a Decision Case
|
|
5. Follow the entire chain from observation to decision with full traceability
|
|
|
|
**When this works, you have the first complete, verifiable Control Intelligence pipeline.**
|
|
|
|
## Development Phases
|
|
|
|
### Phase 1 — Essential (for pilot missions)
|
|
|
|
| Feature | Purpose |
|
|
|---------|---------|
|
|
| **Ingestion** | Upload images/video, show metadata (GPS, time, device), store raw data immutably |
|
|
| **Dataset Explorer** | Browse missions, filter, search, open a mission |
|
|
| **Annotation** | AI suggestions, manual correction, version history |
|
|
| **Decision Case Viewer** | Observation → Evidence → Finding → Decision → Business Impact → Learning |
|
|
|
|
**This is the heart.**
|
|
|
|
### Phase 2 — Scale (when running many missions)
|
|
|
|
| Feature | Purpose |
|
|
|---------|---------|
|
|
| **Replay** | Compare model versions on same mission |
|
|
| **Benchmark** | Model comparisons |
|
|
| **Evaluation** | Regression tests |
|
|
|
|
### Phase 3 — Advanced (when having multiple models)
|
|
|
|
| Feature | Purpose |
|
|
|---------|---------|
|
|
| **GPU Jobs** | Training queue |
|
|
| **Hyperparameter Runs** | Automated experiments |
|
|
| **Model Promotion** | Development → Validation → Pilot → Production |
|
|
| **Canary Releases** | Gradual rollout |
|
|
|
|
## New Areas
|
|
|
|
### Data Quality
|
|
|
|
Before training anything:
|
|
|
|
```
|
|
Images
|
|
├── Healthy
|
|
├── Blurred
|
|
├── Duplicate
|
|
├── Wrong GPS
|
|
├── Night
|
|
├── Rain
|
|
└── Occluded
|
|
```
|
|
|
|
Coverage:
|
|
```
|
|
Roads, Buildings, Signs, Drainage, Vegetation
|
|
```
|
|
|
|
Quickly shows why a model performs well or poorly.
|
|
|
|
### Decision Analytics
|
|
|
|
Not AI metrics. Decision metrics.
|
|
|
|
| Metric | Question |
|
|
|--------|----------|
|
|
| Acceptance Rate | How many recommendations were accepted? |
|
|
| Ignore Rate | How many were ignored? |
|
|
| Accuracy | How many proved correct? |
|
|
| Insufficient Evidence | How many lacked enough evidence? |
|
|
| Data Collection Value | How often did "Collect more data" lead to a better decision? |
|
|
|
|
Closer to LandveX business value than traditional AI metrics.
|
|
|
|
---
|
|
|
|
## ändringshistoria
|
|
|
|
| Version | Datum | Beskrivning |
|
|
|---------|-------|-------------|
|
|
| 1.0 | 2026-07-02 | Initial specification for LandveX Intelligence Lab |
|
|
| 1.1 | 2026-07-02 | Added architecture, pipelines, immutability |
|
|
| 1.2 | 2026-07-02 | Added MVP milestone, development phases, Data Quality, Decision Analytics |
|
|
|
|
---
|
|
|
|
## STATUS
|
|
|
|
**SPECIFICATION — Awaiting development decision**
|