Skip to content

Streaming uploads and downloads across storage clients and storages #2240

Description

@vdusek

apify-client-python is adding streaming request bodies in apify/apify-client-python#1060. After it lands, set_record accepts a file-like object, an iterator of bytes/str chunks, or a streamed HttpResponse, and sends it chunked without buffering. The download direction (stream_record) has been there for a long time.

Crawlee has no equivalent at any layer: not in the KeyValueStoreClient / DatasetClient contracts, not in the KeyValueStore / Dataset frontends, and not in the backends. Passing a file object through today's API also corrupts the record silently on three of the five KVS backends.

#1931 asks for streaming KVS records and is still in solutioning. This issue is the wider investigation it needs: what the interfaces should look like, what each backend can actually do, and what the storages should expose.

Related

✍️ Drafted by Claude Code

Activity

  1. added
    t-toolingIssues with this label are in the ownership of the tooling team.
    on Sep 16, 2026
  2. vdusek commented on Sep 16, 2026

    @vdusek
    CollaboratorAuthor

    What happens today

    Uploads

    KeyValueStore.set_value (src/crawlee/storages/_key_value_store.py:173) forwards value: Any to KeyValueStoreClient.set_value (src/crawlee/storage_clients/_base/_key_value_store_client.py:53). Neither says what value may be, and the backends disagree:

    • File system (_file_system/_key_value_store_client.py:302), SQL (_sql/_key_value_store_client.py:149) and Redis (_redis/_key_value_store_client.py:125) serialize through a chain that ends in str(value).encode('utf-8'). A file handle or a generator is stored as its repr, so await kvs.set_value('out', open('big.bin', 'rb')) writes the literal text <_io.BufferedReader name='big.bin'> and reports success.
    • Memory (_memory/_key_value_store_client.py:112) keeps the object as it is and records sys.getsizeof(value) as the size. A generator is then consumed by whichever reader gets to the record first.
    • Apify (apify-sdk-python, src/apify/storage_clients/_apify/_key_value_store_client.py:123) hands the value straight to set_record, so once feat: Stream request bodies from files, iterables, and responses apify-client-python#1060 lands this one backend streams correctly by accident, with nothing in Crawlee's types to promise it.

    So the same call is a working streaming upload on the platform and silent data loss on the local default.

    Downloads

    Every backend materializes the whole record. The file system client does record_path.read_bytes(), SQL selects the full BLOB, Redis does hget, and the Apify client calls get_record even though stream_record is available. KeyValueStore.get_value returns a decoded value and there is no way to ask for a stream, so reading a 2 GB screenshot or a large NDJSON export costs 2 GB of RSS.

    Datasets

    Dataset.export_to (src/crawlee/storages/_dataset.py:362) collects every item into a list (src/crawlee/_utils/file.py:193), serializes it into a StringIO, then passes the resulting str to set_value. A large dataset is held in memory at least twice before the first byte is written. This is the most obvious consumer of a streaming set_value, since the export is already produced incrementally.

    On the read side, Dataset.iterate_items already pages, so the gap there is smaller. Worth checking whether get_data with the default limit=999_999_999_999 deserves the same treatment.

    Questions to settle

    1. Where does the contract live? Does KeyValueStoreClient.set_value widen its value type to include file-likes and (async) iterators, or does streaming get its own method (stream_value / set_value_stream) that backends opt into? The first keeps one call site, the second keeps Any from meaning "anything, results may vary".
    2. What do non-streaming backends do? SQL and Redis store one value per row/field, so they cannot stream a write in any useful sense. Buffering the source into bytes is correct and at least well defined, unlike today's repr. The file system backend can stream a chunked write for real.
    3. What shape does the download take? An AsyncIterator[bytes], an async context manager wrapping the response, or a file-like object. It has to work for a local file, a Redis field, and an HTTP response, and it has to make the Apify backend's stream_record usable through the storage.
    4. What happens to the record metadata? infer_mime_type cannot inspect a stream, so the content type has to default to application/octet-stream or be required from the caller, and size is unknown until the source is drained.
    5. Retries and seekability. apify-client-python retries only seekable sources, rewinding before each attempt, and gives everything else a single attempt. Crawlee's backends need a rule of their own for a half-written file or a partially consumed iterator.
    6. Should Dataset.export_to stream into the KVS once set_value supports it? That removes the StringIO and the fully materialized item list.
    7. Does anything on the JS side need to match? KeyValueStore.setValue there already accepts a stream, while reading one means dropping down to the resource client (Better KVS streaming support crawlee#2929).

    ✍️ Drafted by Claude Code

  3. added
    solutioningThe issue is not being implemented but only analyzed and planned.
    on Sep 16, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    solutioningThe issue is not being implemented but only analyzed and planned.t-toolingIssues with this label are in the ownership of the tooling team.

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions