Skip to content

ML-Based Intelligent Data Tiering

UVP

Most three-tier storage systems force you to write rules: “move data older than 30 days to cold”. HeliosDB Full’s tiering layer learns the access pattern itself. An ML ensemble predicts whether each object should live on Hot NVMe, Warm SATA-SSD, or Cold S3, then a cost-aware optimizer migrates only when the move is net-positive against your latency SLA. Production deployments report 60-85% storage cost reduction with <10% p95 latency impact and 82-87% prediction accuracy.


Prerequisites

  • HeliosDB Full v8.0.3 or later
  • At least three storage tiers configured (NVMe, SATA-SSD, object store — or any subset)
  • A workload with at least 100 access samples (the ML model needs training data; below that, the rule-based fallback kicks in)
  • Prometheus or any metric scraper (recommended; the tiering layer exports rich metrics)
  • ~25 minutes

ML tiering sits on top of the base 3-tier storage manager — this tutorial assumes the base tiering is already configured.


1. The Three Tiers

TierBackingLatency targetCost / GB / moUse case
HotNVMe SSD~1 ms$0.15Active rows, high IOPS
WarmSATA SSD~5 ms$0.04Moderate access, working set
ColdS3 Standard / Azure Blob / GCS~50 ms$0.02Archive, infrequent reads

These figures are the defaults; override them in the tiering configuration to match your actual cloud bill. The optimizer uses the cost numbers you give it — if they’re wrong, the migration plan is wrong.


2. The Background Loop

The ML loop runs every 5 minutes (background_job_interval_secs: 300), retrains every 24 hours, and only migrates when confidence ≥ 0.75 and the projected savings beat the migration cost.


3. The Six Components

┌─────────────────────────────────────────────┐
│ ML Tiering Orchestrator │
└─────────────────────────────────────────────┘
│
┌────┼────┐
▼ ▼ ▼
┌────────┐ ┌──────────┐ ┌────────────┐
│Predictor│→│Optimizer │←│Policy │
│ (ML) │ │ (Cost) │ │Engine │
└────────┘ └──────────┘ └────────────┘
▲ │ │
┌────────┐ ┌──────────┐ │
│Monitor │ │Migrator │←───────┘
└────────┘ └──────────┘
ModuleJob
Access Pattern PredictorML ensemble, confidence-scored
Cost ModelPer-tier $ math, savings projections
Tier OptimizerCost-aware ranking, latency-bounded plan
Policy EnginePin / exclude / threshold rules — overrides ML
Access MonitorRecords every get/put for the trainer
Data MigratorBandwidth-capped, retry-with-backoff movement

The policy engine takes precedence over the ML model — useful when the model is still learning, or when compliance forces a tier (e.g. EU customer data must stay on EU-region warm storage).


4. How the Optimizer Decides

The optimizer migrates an object only when the projected savings outweigh the expected latency penalty and the migration cost, and the model’s confidence meets confidence_threshold. That second clause is the safety net — even if the math says “move it”, the model has to be confident enough.

Worked example — a 1 TB deployment

StrategyHotWarmCold$ / monthAnnual
Baseline (all hot)100%——$150$1,800
ML-optimized5%25%70%$31.50$378

Saving: $1,422 / year (79% reduction) for a 1 TB workload. Multiply by 100 for a 100 TB enterprise — that’s $142,200 / year off the cloud bill from a single feature flag.


5. Policy Engine — When the ML Model Is Wrong

The model is wrong sometimes. Real-world cases:

  • New product launches: no access history yet, the model is cold-starting
  • Compliance pin: “this customer’s data must live on-prem”
  • Predictable bursts: end-of-month billing reads

Three rule types cover these: pin rules (hard-pin a path to a tier, overriding ML), exclude rules (e.g. compliance — never move a path to cold), and threshold rules (e.g. promote anything accessed more than N times a day).

The policy engine runs before the optimizer. Pin rules win; the ML model only decides among the tiers the policy hasn’t already constrained.


6. Configuration

Recommended starting points:

  • Tighten confidence_threshold to 0.85 in the first 7 days; relax to 0.75 once the model converges
  • Cap max_concurrent_migrations at 10% of your I/O budget
  • Set max_latency_degradation based on your customer-facing SLA — the optimizer treats it as a hard constraint

7. Observability

Prometheus metrics exported:

heliosdb_tiering_predictions_total{model="ensemble",outcome="correct|wrong"}
heliosdb_tiering_migrations_total{from_tier,to_tier,result}
heliosdb_tiering_bytes_moved_total{from_tier,to_tier}
heliosdb_tiering_savings_dollars{period="monthly"}
heliosdb_tiering_latency_p95_seconds{tier}

8. Performance Reference

Indicative figures measured on representative fixtures; reproduce on your own hardware.

OperationLatencyThroughput
Feature extraction (1000 objects)47 ms21,277 obj/s
Cost calculation (1000 ops)0.8 ms1,250,000 ops/s
ML training (100 samples)950 ms—
ML training (1000 samples)5.4 s—
Policy application (1 rule)12 µs—
Access recording (1000 ops)120 ms8,333 ops/s

The 24-hour retraining cycle is the dominant background cost; on a 1000-sample workload it takes ~5 seconds and runs once per day.


9. Production Checklist

  • Three tiers are configured, with realistic cost_per_gb numbers from your provider invoice
  • At least 100 access samples have been collected before enabling auto-migration (use auto_tiering_enabled: false while warming up)
  • max_latency_degradation reflects a real SLA, not a guess
  • Pin rules cover the obvious “never move this” categories (billing, compliance, audit logs)
  • Bandwidth limit is set to a sane fraction of total I/O (10% is a good first guess)
  • Prometheus is scraping; an alert exists on heliosdb_tiering_predictions_total{outcome="wrong"} rate
  • The cold tier object store is the same region as the rest of your stack (egress is the silent killer)

Where Next

  • Cognitive Agents — pair the storage tiering loop with the schema-manager agent for full-stack autonomy
  • PITR Recovery — make sure cold-tier objects participate in your PITR plan
  • Multi-Tenancy Setup — per-tenant tiering policies via the policy engine