🎀 Reverie, the Summit for AI Builders

Every generative model is a kind of dream, a plausible world rendered from what it has seen. Reverie, the summit for AIΒ builders by LanceDB, is taking place on November 5 in San Francisco. Researchers and engineers behind frontier models, world models, generative video, and physical AI go deep on the data systems those dreams are made on.

Featuring speakers from NVIDIA, Runway, Luma, Applied Intuition, Cruise, Adobe, Exa, and more.

Apply to Attend β†’ | Read more β†’

🧠 Why CrewAI Rebuilt Agent Memory on LanceDB, Powering 2B+ Agent Executions

CrewAI replaced a two-system memory stack with a single LanceDB table storing vectors, metadata, and multimodal data together, scoring recall by similarity, recency, and importance so old critical decisions still surface over trivial recent ones. Now CrewAI's default vector backend, it powers 12M monthly downloads and 2B+ agent executions.

Read more β†’

⚑ Data Loading for AI/ML: A Comprehensive Guide

GPU training throughput rarely bottlenecks on I/O β€” the real trap is the CPU stage, where JPEG decoding and tokenization can starve even a capable pipeline. Weston Pace walks through the three-stage model with concrete numbers (16K tokens/sec on an L40S, 1GB/s per core) and practical PyTorch + LanceDB guidance.

Read more β†’

πŸ“š Also Published

πŸ“ΈΒ Actuate & Ray Summit in SF

LanceDB team at Actuate and Ray Summit this month! Thanks to everyone who stopped by our booth and said hi. We really enjoyed meeting so many of you in person and talking about data infrastructure for physical AI, robotics, and foundation models. See you at our own Reverie on November 5 πŸ‘‹

πŸ“… Upcoming Events

Reverie β€” November 5, 2026 Β· San Francisco, CA

Reverie is a one-day technical summit for researchers and engineers building the data systems behind world models, generative video, physical AI, multimodal search, and agentic research.

Apply to Attend β†’

πŸ—οΈ LanceDB Enterprise Updates

Performance

  • Less job-history scanning on cleanup β€” Per-job scratch storage removes a cleanup job's need to scan every other job's record for what's expired; on a 604-record benchmark, that scan alone cost ~600 sequential storage reads and 15.2 of 26.3 total seconds (58%).
  • Fewer redundant manifest reads on WAL nodes β€” Reusing a shard's own manifest cache instead of rebuilding it each read cuts a flush operation's storage requests from 132 to 12, and stops idle tables from burning ~17 requests/minute on background GC and compaction checks with no client traffic.
  • Concurrent SSTable fetch in WAL compaction β€” Fetching a compaction pass's SSTables concurrently instead of one at a time cuts fetch time by roughly 4x at the default fan-out: an 8-SSTable pass drops from 1.6s to 0.4s, and a 32-SSTable pass from 6.4s to 1.6s.
  • Bounded admin traffic on large clusters β€” Capping concurrency on cluster-status requests and polling only the replica-group members a job actually targets, instead of every member cluster-wide, prevents a connection spike on the query node as replica-group counts grow.
  • Streaming computed-column refresh β€” Streaming a computed column's output directly into its fragment file, instead of buffering the whole output twice, removes the artificial cap that previously bounded how much a single refresh batch could compute at once.

Features

Feature Description
Materialized views SQL now supports CREATE, REFRESH, and SHOW MATERIALIZED VIEW, a filtered or projected view of another table that refreshes as a background job, with a validation token guarding against refreshing a view dropped and recreated under the same name.
SQL-defined computed columns SQL now supports registering Python functions as computed columns, with WHERE-filtered refresh to recompute only matching rows and GPU-backed execution for functions that need it, bringing Feature Engineering's computed-column workflow into SQL.
Table cloning in SQL CREATE TABLE ... CLONE now works from the SQL shell and Flight SQL, closing the last surface that lacked it: clone a table as of its latest version, a specific version, a tag, or a timestamp, without copying the underlying data.
Column lineage graph A new lineage view shows which upstream columns and Feature Engineering jobs produced each column, rendered as an interactive graph with search and click-to-pin, keeping a complex table's derivation history legible instead of an unreadable tangle.
Row and fragment sampling The SQL:2003 TABLESAMPLE clause is now supported for both row-level and whole-fragment sampling, by percentage or exact row count, with a repeatable seed for reproducible samples and support for sampling one side of a join independently.

‍

🌟 Open Source Releases

Project Description
Lance v3.0.2 – v11.0.0
Release notes
β€’ ACORN-1 traversal cuts worst-case HNSW search latency by 10–250x when a filter clusters away from the query's region, at up to 4 points of recall (opt-in via ApproxMode::Fast) (#7927); a new MAXSCORE algorithm for pure-SHOULD FTS queries cuts candidate probes by up to 53.6x in benchmark testing (#8474); and FTS scorers can now compose across columns in one shared row-address domain (#8685)
β€’ Segmented indexes now support ngram (#7244), bloom filters (#7925), and RTree (#7932); zone maps extended to all data types (#8017) with seeds written into data file footers (#7427)
β€’ New Dataset.migrate_to_stable_row_ids migration method (#8521), cross-store deep_clone support (#7545), compaction limits via max_source_rows/max_source_bytes (#8235), and per-fragment column writes that survive compaction (#8313)
LanceDB v0.37.1 – v0.38.0
Release notes
β€’ Materialized views land in the client SDKs β€” declare and refresh on local tables (#3930, #4010) with Python (#3933) and Node (#3935) bindings, alongside SQL-declared computed columns (#3937) with async refresh returning a job handle (#3939)
β€’ New Function framework for registering scalar UDFs with typed wire contracts (#3985), grouped column bindings (#3994), GPU resource requirements (#4085), and Blob v2 signatures (#4091)
β€’ Streaming dataset adds backpressure on its post-transform queue (#3897) and sequence packing (#3920); the data loader now reads remote tables (#3981) pinned to a fixed base table version (#3982)
lance-namespace v0.11.0 – v0.11.1
Release notes
β€’ InsertIntoTableResponse now returns num_inserted_rows and version fields (#359)
β€’ Support for declaring and backfilling expression-computed columns (#360)
β€’ Vector index build parameters (e.g. IVF/PQ settings) can now be specified via CreateTableIndexRequest (#361)
lance-trino v0.3.4
Release notes
β€’ Fixed Substrait name emission for structs nested inside lists, improving compatibility with complex schema queries (#227)
β€’ Optimized COUNT() queries over non-null constants (#155)

🫢 Community Contributions

Thank you to contributors from Netflix, Bytedance, Adobe, Intel, Microsoft, Pinterest, Huawei for improvements across storage, indexing, query execution, distributed processing, and ecosystem integrations in LanceDB, Lance, and the broader ecosystem.

Notable contributions this month:

  • @yanghua β€” Added a pluggable cache backend API to the Java SDK, letting custom caching strategies register and switch at runtime on top of Lance's existing cache registry
  • @zhangyue19921010 β€” Introduced compaction controls with max_source_rows and max_source_bytes limits, enabling fine-grained resource management during file compaction
  • @sezruby β€” Exposed ArrowArrayStream export on LanceScanner in Java, enabling zero-copy interoperability with Arrow-native tooling
  • @morales-t-netflix β€” Added per-value support for fixed-length packed structs, enabling efficient columnar encoding for complex nested types
  • @leohoare β€” Implemented ACORN-1 traversal for prefiltered HNSW search, cutting worst-case query latency 10–250x when a filter clusters away from the query's region, at up to 4 points of recall (opt-in via ApproxMode::Fast)
  • @xtangxtang β€” Fixed fp16 IVF partition assignment to route through AMX-FP16 GEMM, recovering misassigned vectors and accelerating vector indexing on Intel hardware
  • @professor-moody β€” Contributed critical security fix bounds-checking length prefix in VariableFullZipDecoder, preventing potential memory safety issues
  • @XuQianJin-Stars β€” Implemented safe commit with ConditionalPutCommitHandler for GooseFS, enabling consistent writes on distributed filesystems
  • @wombatu-kun β€” Fixed stale per-segment rows in vector search after in-place column value updates, keeping the index consistent with the latest data

A heartfelt thank you to our community contributors of Lance and LanceDB this past month:

@1fanwang β€’ @3286360470 β€’ @a-erofeev β€’ @adityaj0 β€’ @ali2arslan β€’ @anirudh-s-kumar β€’ @anonx3247 β€’ @antio2 β€’ @ar-maan05 β€’ @arielyes β€’ @atirna β€’ @beinan β€’ @brunosrz β€’ @clearlove10-c β€’ @com-junkawasaki β€’ @cswpy β€’ @dawid0309 β€’ @dcfocus β€’ @ddupg β€’ @dentiny β€’ @divyanshus2404 β€’ @dshepelev15 β€’ @dtolnay β€’ @dubin555 β€’ @ecthlion β€’ @everysympathy β€’ @fangbo β€’ @fanng1 β€’ @farazshaikh β€’ @fzowl β€’ @geserdugarov β€’ @haochengliu β€’ @hellower β€’ @hfutatzhanghb β€’ @huahuay β€’ @igorganapolsky β€’ @ilya-zlobintsev β€’ @isaac-dasari β€’ @ivscheianu β€’ @j7nhai β€’ @jackylee-ch β€’ @jagrutipatilp β€’ @jay-ju β€’ @jayson-huang β€’ @jerryjch β€’ @jiaoew1991 β€’ @jiaqizho β€’ @jimmy-xie-fleet β€’ @jo-migo β€’ @jonasdedden β€’ @jsnider3 β€’ @julianyg β€’ @kamronis β€’ @kangnan-li β€’ @keunhong β€’ @leoreeyang β€’ @lh-kevin β€’ @lichuang β€’ @luciferyang β€’ @lucyge2022 β€’ @majin1102 β€’ @mannxo β€’ @maswin β€’ @medisean β€’ @mikemikimike β€’ @mikewhb β€’ @mmatczuk β€’ @nyl3532016 β€’ @paramt β€’ @pengw0048 β€’ @pjdurden β€’ @primorlee β€’ @prrao87 β€’ @puchengy β€’ @qiuyuhang β€’ @ragingkore β€’ @ragnorc β€’ @raygao25 β€’ @ringfa11 β€’ @risto0211 β€’ @roridemonslayer β€’ @rupertmaiti2005 β€’ @sbrunk β€’ @seven7763 β€’ @sliortega295-ops β€’ @sravan1011 β€’ @stevestevenpoor β€’ @tandede β€’ @touch-of-grey β€’ @u70b3 β€’ @valkum β€’ @vatharevinayak β€’ @vinaysurtani β€’ @vip892766gma β€’ @weimingdiit β€’ @winklemad β€’ @wirybeaver β€’ @xiaguanglei β€’ @xixigoodluck β€’ @xloya β€’ @yentur β€’ @yichenw β€’ @yuvalif β€’ @yuw1 β€’ @zackfairts β€’ @zhangstar333 β€’ @zouhuajian β€’ @zyt-yt

🀝 Lance Community Sync Recap

The community syncs this month covered the Lance 10.0.0 SDK release, which included a critical encoding fix backported to all 0.x versions, along with a refactor of file version handling. Discussion also focused on the Lance 11.0.0 SDK beta development and its progression to RC2, with work done to address benchmark regressions before release. Process changes were announced for format-change votes, moving from GitHub discussions to pull requests with a shortened 72-hour window. The team also discussed proposals for experimental feature flags and native video/blob data types in Lance.

The next Lance Community Sync will take place on Thursday, September 10 @ 9am PT.

‍