Platform
The technology behind the research.
Data, models and risk connected through a shared research and engineering framework.
The platform brings market data, financial models and distributed computing into one research environment. Shared definitions and traceable outputs connect the work, from constructing a futures history to examining a portfolio’s exposures.
A connected research platform
Select an area to explore its methods.
Risk estimation runs alongside model research; forecasts and risk meet in portfolio analysis.
How it fits together
A reusable framework.
A focused research implementation.
Aspen provides the common building blocks. Savannah defines the markets, features and research configurations that use them. Shared interfaces allow an experiment to change without rebuilding the surrounding data and execution machinery.
Aspen
Data interfaces, instrument objects, feature computation, training, risk and portfolio algorithms.
Explore the framework designSavannah
Market universes, curve and feature definitions, model configurations and research workflows.
Explore research definitionsExplore the platform
Six areas, from the architecture to the methods and implementation behind it.
Data & Provenance
Typed datasets, traceable revisions and explicit publication controls turn vendor inputs into a consistent research foundation, with the information needed to inspect data quality and reproduce a run.
Explore 5 areas
Can the research inputs be trusted, traced and replayed?
Market data acquisition
- Vendor adapters and Barchart acquisition
- Daily and intraday prices, volumes and open interest
- ETF, index, rate and security metadata
- Raw snapshots, input manifests and replay
Schemas and analytical storage
- Typed dataset and metadata schemas
- Partitioned Parquet on local disk and S3
- DuckDB SQL, date bounds and partition pruning
- Separation of market data, features, research and temporary outputs
Futures history construction
- Contract discovery, expiries and vendor exceptions
- Continuous contracts and roll policies
- Adjusted and unadjusted histories
- Curve-ready futures datasets
Quality, revisions and lineage
- Checks, quarantine and controlled repair
- Run identity, provenance and validated data reads
- Revisions, rebuilds and snapshot identity
- Observation time, availability time and historical-vintage limits
Controlled publication and efficient reads
- Stage, validate, promote and verify
- Metadata-gated visibility
- Publication indexes, change markers and safe fallback
- Cache locality, bounded reads and dataset readiness
Instruments, Curves & Features
Consistent definitions of instruments, curves and relative-value relationships connect market structure to a configurable feature library for model and portfolio research.
Explore 5 areas
How are market structure and investment ideas represented?
Market representations
- Outrights: futures, ETFs, indexes and rates
- Asset return residuals after controlling for multi-factor risk exposures
Curves and term structures
- Ordinal contract-rank and time-to-maturity coordinates
- Futures and rate-curve representations
- Polynomial, spline and Nelson–Siegel fitting
- Fit parameters, residuals and quality diagnostics
Relative value instruments
- Calendar and cross-market spreads
- Butterflies and relative curve shape
- Raw, volatility-scaled and beta-hedged definitions
- Hedge estimation, return semantics and tradeable-leg mapping
Feature language and library
- Declarative functions, inputs, arguments and composition
- Curve slope, curvature, dispersion and relative structure
- Spread/fly features and underlying-leg context
- Technical, volatility, liquidity and open-interest features
- Calendar, intraday and transaction-cost predictors
Feature processing and execution
- In-memory feature objects and shared instrument state
- Validation, warm-up, normalization and seasonality
- Asset-centric build planning and cache reuse
- Ray/loky execution and incremental publication
- Adding configured variants versus adding a new transformation
Modelling & Backtesting
Shared data and model interfaces support rolling experiments across targets, horizons and model families, with explicit timing, reusable computation and diagnostics that make research results easier to examine.
Explore 6 areas
Can a hypothesis be tested consistently across time, markets and models?
Targets and temporal contracts
- Full return, regression residual and factor-adjusted targets
- Forecast origin, execution lag and outcome maturity
- Outright, spread, curve and butterfly targets
- Settlement-window range and volume targets
- Training-data interfaces and inference alignment
Feature universes and reduction
- Global, asset-class and local feature groups
- IC ranking and point-in-time feature selection
- PCA, PLS and reusable fitted reductions
- Scaling, numerical guards and target inverse transforms
Extensible model families
- Ridge, Lasso and ElasticNet regression
- Logistic and tree classification interfaces
- Random forest, XGBoost and LightGBM
- Common model adapters and configurable grids
- Optional SHAP and model diagnostics
Rolling and distributed experiments
- Walk-forward training and daily/incremental modes
- Research backtests versus portfolio backtests
- Deterministic plans, training-run identities and partitioning
- Ray workers, bounded loky work and thread budgets
- Shared data and reduction caches, saved model runs and resumability
Backtesting and simulation
- Walk-forward training and out-of-sample forecast evaluation
- Historical portfolio backtests with evolving positions and transaction costs
- Holding-period, turnover and trading-cost sensitivity
- Monte Carlo simulation using block-bootstrap return paths
- Risk-model scenario analysis and portfolio sensitivity
- Drawdown, margin and capital requirements under stress
Forecast evaluation and research outputs
- Predictive strength and calibration: Pearson/Spearman IC, directional accuracy, calibration slopes and forecast scale
- Stability and breadth across periods, feature groups, training windows and model parameters
- Horizon efficacy: frozen-holding payoffs, positive outcome share and selection coverage
- Portfolio outcomes: cumulative returns, Sharpe ratios, realised volatility and drawdowns
- Implementation sensitivity: turnover, trading costs, signal smoothing and long/short exposures
- Matched-date baseline comparisons and block-bootstrap confidence intervals
- Interactive diagnostics: forecast curves, holdings, portfolio weights and downloadable results
Risk & Portfolio Construction
Factor models, portfolio optimisation and scenario analysis connect forecasts to exposures, costs and capital requirements, providing tools to examine the choices between a statistical signal and a practical position.
Explore 6 areas
How do forecasts connect to exposures, portfolio shape and capital requirements?
Multi-factor risk modelling
- Economic factor definitions and ordered orthogonalization
- Latent PCA components
- Rolling ridge exposures, intercepts and residuals
- Covariance, specific risk, shrinkage and snapshots
Portfolio geometry and allocation
- Expected return, variance and risk decomposition
- Factor concentration, diversification and effective breadth
- Mean-variance and hierarchical optimization
- Objectives, constraints, stage retention and solver fallback
Scenarios and capital
- Historical and block-bootstrap return paths
- Historical risk-state selection and deterministic perturbations
- Drawdown, losses, margin and survival-capital estimates
- Leverage capacity and shrink-only risk overlays
Risk diagnostics and adaptive exposure
- Fast versus reference risk-model disagreement
- Realized versus modelled volatility and diversification
- Optimizer sensitivity to risk perturbations
- Calibration, smoothing and exposure attenuation
Costs and implementation
- Spread-cost estimates and cost stress
- Forecast range and volume as liquidity inputs
- Turnover penalties and position changes
- Exposure weights, contract units and margin semantics
- Position sizing and execution integration boundaries
Portfolio validation
- Signal backtests versus portfolio P&L
- Full-return attribution and factor exposure
- Cost, liquidity and execution sensitivity
- Risk-model calibration, scenario sensitivity and integration tests
Cloud Infrastructure & Operations
Workload-specific cloud compute runs research jobs alongside persistent storage and orchestration, with pinned releases, progress monitoring and recovery procedures that keep large experiments traceable and manageable.
Explore 5 areas
How does the platform run at scale without a permanently large compute estate?
System topology
- Windows acquisition/control VM and Linux compute
- S3 storage boundaries and immutable artifacts
- EC2/Ray head and worker roles
- Dedicated raw, futures, feature, learning and risk jobs
- SSM access, workload IAM and S3 read/write boundaries
Research compute capacity
- Infrastructure templates and environment configuration
- Worker sizing, quotas and bounded concurrency
- Provisioning configured capacity, run completion and cleanup
- Data locality, startup overhead, resource budgets and cost measurement
Build and release
- Paired Aspen/Savannah source revisions
- Docker build and smoke validation
- ECR digest promotion and release pointers
- SSM resolution and run-pinned container images
- Development/production boundaries and rollback
Runtime orchestration
- Daily, seed, incremental and historical research jobs
- PowerShell/operator launchers and remote entrypoints
- Research destination isolation
- Readiness checks, timeouts and lifecycle controls
Monitoring and recovery
- CloudWatch logs and progress inventories
- Dashboard refresh, freshness and failure visibility
- Planned, active, cached, published and failed work
- Publication recovery without unnecessary retraining
- Operator runbooks and incident-driven improvements
Engineering & Reliability
Explicit interfaces, versioned configurations and focused tests make the platform extensible and reviewable, while failure handling and documentation help research components work together as a coherent system.
Explore 5 areas
What makes the work maintainable, extensible and reviewable?
Framework architecture
- Aspen core contracts and reusable algorithms
- Savannah research definitions and operating policy
- Configuration, factories and extension points
- Dependency direction and paired-version compatibility
Reproducibility
- Configuration fingerprints and experiment identity
- Code, image, dataset and run provenance
- Deterministic planning and consistent storage contracts
- Recreating a result versus historical information-set fidelity
Testing and assurance
- Unit and numerical-invariant tests
- Schema, temporal-alignment and configuration tests
- Integration, failure/recovery and environment tests
- Continuous integration design and test evidence
Failure design
- Idempotent attempts and bounded retries
- Quarantine, partial outputs and publication barriers
- Cache recovery and explicit fallback paths
- Resource and memory bounds
Documentation and implementation evidence
- Architecture decisions and module documentation
- Operating runbooks and reproducible examples
- Framework design and infrastructure examples
- Versioned evidence links and curated code tours
- Research records and publication boundaries