Aerospike
Description
The Aerospike online store provides support for materializing feature values into an Aerospike cluster for serving online features.
The Aerospike online store is currently in preview. Some functionality may be unstable, and breaking changes may occur in future releases.
Features
Supports both synchronous and asynchronous read/write paths (
online_read/online_read_async,online_write_batch/online_write_batch_async). Async methods wrap the blocking client inrun_in_executor, keeping the event loop responsive in feature-server workloads.Partial, server-side upserts via Aerospike Map CDT operations — writing one feature view never clobbers another feature view stored on the same entity.
Record-level TTL controlled by a single
ttl_secondsconfig option (honours the namespace default, a "never expire" sentinel, or an explicit number of seconds).Per-feature-view namespace overrides and set overrides — pin individual feature views to RAM-only or SSD-backed namespaces, or isolate one view in its own set, without splitting projects.
Prewriting hook — a configurable, import-string-resolved callable applied to every write batch for cross-cutting concerns like PII masking, application-side encryption, or value coercion.
Authentication and TLS options for Aerospike Enterprise Edition passed straight through to the Aerospike Python client.
client_kwargsescape hatch for any advanced client-config field not surfaced onAerospikeOnlineStoreConfig.Baseline: Aerospike Server ≥ 6.0 (uses batch-write / batch-operate APIs). The store has been developed against CE 8.x.
Getting started
Install the Aerospike extra (alongside the dependency for the offline store of choice):
You can start from any of the standard templates (e.g. feast init -t local or feast init -t aws) and then swap in Aerospike as the online store as shown below.
Examples
Basic configuration — local Aerospike CE
Multi-node cluster
Timeout semantics. The Aerospike client distinguishes per-attempt (
socket_timeout) from total (total_timeout) deadlines.*_timeout_msmap tototal_timeout— the overall budget for a call including retries. Setsocket_timeout_msas well so each individual attempt has its own (shorter) deadline; without it,max_retrieseffectively never fires because the first attempt is allowed to consume the entire total deadline.
Batch chunking.
online_readandonline_write_batchsplit large requests into chunks of at mostbatch_max_records(default1000). Aerospike enforces a per-node batch limit via the serverbatch-max-requestssetting (historically5000). Lowerbatch_max_recordsif your cluster cap is tighter; raise it only when the server limit and client timeouts allow.
Aerospike Enterprise with authentication
Requires Aerospike Enterprise Edition. The Community Edition server has no built-in user/security model and will reject these config keys.
Aerospike Enterprise with TLS
Requires Aerospike Enterprise Edition. The Community Edition server does not implement TLS, so
tlsconfig is effective only against EE clusters.
Per-feature-view namespace and set overrides
Two Dict[str, str] config fields — namespace_overrides and set_overrides — let you place individual feature views on a different Aerospike namespace or set without splitting your project across stores. Anything not listed in either map falls back to the store-level default (namespace / set_name_template).
Common reasons to reach for these:
A hot, latency-sensitive view belongs on a RAM-only namespace; a wide, cold view belongs on an SSD-backed namespace. Same project, different storage tiers.
You want
feast applydeletions ortruncateon one feature view to be O(1) without scanning records of the others — give that view its own set.
Tradeoffs.
Every namespace listed in
namespace_overridesMUST already exist on the cluster — Aerospike cannot create namespaces at runtime, and a missing namespace surfaces as an opaqueAEROSPIKE_ERR_PARAMon the first read or write.Putting feature views on different sets means a multi-feature-view read for the same entity becomes one Aerospike round trip per set, not one round trip total. Only opt in when the operational isolation is worth that cost. Reads that touch a single feature view are unaffected.
Admin operations honour the overrides automatically:
update()(called byfeast apply) groups dropped feature views by their resolved(namespace, set)and issues one background scan per group;teardown()truncates every unique(namespace, set)pair the project may have written to (including the store-level default).
Prewriting hooks
prewriting_hook is the import path of a callable that is invoked once per online_write_batch call, receives the rows about to be written, and returns the rows that actually go on the wire. Use it for cross-cutting write-side concerns that you don't want sprinkled through every materialization job — PII masking, application-side encryption, dual-write fan-out, value coercion, etc.
Hooks are referenced by import string (rather than as a Python Callable value) so the config survives YAML/JSON serialisation and remote-feature-server transport. The resolved callable is cached on the store instance, so import cost is paid once per store lifetime.
Hook signature:
The hook MUST return a row list with the same schema as its input. Returning [] short-circuits the write — same path as an empty input, no wire call is issued. Hooks that raise will fail the whole batch; there is no per-row fallback.
1. Drop a hook function in your project. Any module on the PYTHONPATH of every process that writes through Feast will do (the materialization workers, the registry CLI host, and the feature server, if you run one).
2. Reference the hook from feature_store.yaml:
Operational notes.
The hook is only invoked on the write path; reads pass through the store untouched. If your hook is one-way (e.g. hashing) you have to apply the same transformation to the candidate value at read time yourself.
Hooks run inside the same process as the writer — they're not RPCs and not sandboxed. They can read environment variables, open files, call out to KMS, etc. Treat them as part of your trusted code base.
A misconfigured
prewriting_hook(bad import path, missing function, non-callable target) raisesValueError/TypeErroron the firstonline_write_batchcall, not on store construction. Add a smoke test that writes one row at deploy time so misconfigurations surface before a real batch.
The full set of configuration options is available in AerospikeOnlineStoreConfig.
Data Model
The Aerospike online store uses a single set per project with entity-key collocation. Features from multiple feature views for the same entity are stored together on a single Aerospike record, analogous to the MongoDB online store's "one document per entity" layout.
Namespace
online_store.namespace (must be pre-configured on the cluster); per-feature-view override via online_store.namespace_overrides
Set
online_store.set_name_template → "{project}_{collection_suffix}" by default; per-feature-view override via online_store.set_overrides
Key
serialize_entity_key(entity_key) as bytearray user key
Bin features
Map CDT keyed by feature-view name, each value a map of feature → native
Bin event_ts
Map CDT keyed by feature-view name, each value an int64 epoch-ms timestamp
Bin created_ts
Top-level int64 epoch-ms timestamp (last feast materialize)
Example record
For a single entity carrying features from two feature views (driver_stats and pricing):
Key design decisions
Record per entity, bin per concept.
featuresandevent_tsare Aerospike Map CDT bins, not dynamic bins, which keeps the store within the 15-byte Aerospike bin-name limit regardless of how many feature views a project has.Partial upserts via Map CDT ops. Writes use
batch_writewithmap_put_items("features", {<fv>: {...}})andmap_put("event_ts", <fv>, <epoch_ms>). Concurrent writes to different feature views on the same entity never clobber each other — each write mutates only its own map keys.Entity-key bytes as the Aerospike user key. Feast's
serialize_entity_keyoutput is passed as abytearrayuser key (notbytes— the Python client hashes only the first byte ofbyteskeys, which would collapse distinct entities).Timestamps as int64 epoch milliseconds. Aerospike has no native datetime type; tz-naive timestamps are treated as UTC per the
OnlineStorecontract.
TTL and expiry
ttl_seconds is written as record-level metadata on every online_write_batch call:
ttl_seconds
Aerospike TTL
Effect
not set / null
TTL_NAMESPACE_DEFAULT
Record inherits the namespace's configured default-ttl.
0
TTL_NEVER_EXPIRE
Record is kept until explicitly deleted.
>0
that many seconds
Record is evicted by the server's nsup thread.
There is no per-feature-view TTL override in this version — the setting is applied uniformly for every write made by the online store.
Indexes
No secondary indexes are created. All access goes through the primary key, which is the serialized entity key.
Async support
Async read/write are provided by running the Aerospike Python client's blocking calls on the default thread-pool executor (loop.run_in_executor). The underlying C client releases the GIL during network I/O, so await store.online_read_async(...) keeps the event loop responsive. A native asyncio Aerospike client is not currently used.
Both sync and async methods are fully supported:
online_read/online_read_asynconline_write_batch/online_write_batch_asyncinitialize/close—initialize(config)eagerly opens the connection so feature servers pay the TCP/handshake cost at startup;close()releases the cached client.
Functionality Matrix
The set of functionality supported by online stores is described in detail here. Below is a matrix indicating which functionality is supported by the Aerospike online store.
write feature values to the online store
yes
read feature values from the online store
yes
update infrastructure (e.g. tables) in the online store
yes
teardown infrastructure (e.g. tables) in the online store
yes
generate a plan of infrastructure changes
no
support for on-demand transforms
yes
readable by Python SDK
yes
readable by Java
no
readable by Go
no
support for entityless feature views
yes
support for concurrent writing to the same key
yes
support for ttl (time to live) at retrieval
yes
support for deleting expired data
yes
collocated by feature view
no
collocated by feature service
no
collocated by entity key
yes
To compare this set of functionality against other online stores, please see the full functionality matrix.
Last updated
Was this helpful?