Skip to content

feat: Add live iteration over a run's dataset items - #1079

Open
vdusek wants to merge 8 commits into
masterfrom
feat/live-dataset-iteration
Open

vdusek wants to merge 8 commits into
masterfrom
feat/live-dataset-iteration

Conversation

@vdusek

@vdusek vdusek commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

Adds RunClient.iterate_dataset_items() and its async twin. They yield a run's dataset items while the run is still pushing them and return once the run has finished and the dataset is drained.

Each poll reads the dataset's itemCount and fetches pages whose limit ends at it. The endpoint scans exactly limit rows, so offsets stay exact even when clean, skip_empty or unwind change how many items a page returns. Between polls, wait_for_finish() waits up to poll_interval for the run to finish.

itemCount lags about 5 s, so after the run finishes, the rows past it are read page by page until a page comes back empty. With clean, skip_empty or unwind, one unfiltered limit=1 read confirms the end. Polling goes on through ABORTING and TIMING-OUT, since such a run can still push items.

The arguments follow DatasetClient.iterate_items without desc and signature, plus poll_interval.

Docs: a new section in the Retrieve Actor data guide, and a mention on the convenience methods and pagination pages.

Closes #1065

✍️ Drafted by Claude Code

@vdusek vdusek added the t-tooling Issues with this label are in the ownership of the tooling team. label Sep 29, 2026
@vdusek vdusek self-assigned this Sep 29, 2026
@codecov

codecov Bot commented Sep 29, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 95.36%. Comparing base (4dcc54c) to head (2254321).
⚠️ Report is 2 commits behind head on master.

Additional details and impacted files
@@            Coverage Diff             @@
##           master    #1079      +/-   ##
==========================================
+ Coverage   95.24%   95.36%   +0.12%     
==========================================
  Files          59       60       +1     
  Lines        5548     5825     +277     
==========================================
+ Hits         5284     5555     +271     
- Misses        264      270       +6     
Flag Coverage Δ
integration 91.51% <94.44%> (-0.54%) ⬇️
unit 87.96% <100.00%> (+0.56%) ⬆️

Flags with carried forward coverage won't be shown. Click here to find out more.

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

@vdusek
vdusek marked this pull request as ready for review September 30, 2026 06:14
@vdusek
vdusek requested a review from szaganek as a code owner September 30, 2026 06:14
@apify-service-account apify-service-account added the tested Temporary label used only programatically for some analytics. label Sep 30, 2026
@vdusek
vdusek requested a review from Pijukatel September 30, 2026 06:35
@vdusek

vdusek commented Sep 30, 2026

Copy link
Copy Markdown
Contributor Author

Hi @barjin and @Pijukatel, FYI: I assigned Jindra to the JS PR and Pepa to the Python one. Both PRs implement the same feature.

Please also let me know WDYT about naming and whether you think adding just one method to RunClient is a good/enough way to support lazy dataset iteration, or whether it should live somewhere else too 🤔.

Thanks!

@Pijukatel Pijukatel left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I wonder if this should be rather implemented inside get_items_iterator as an optional stop_condition_callback(placeholder name), which would be None by default (current behavior) or a user-defined callback that checks whether we should stop or not.

For this specific feature, we would then define this user-defined callback to check if the actor run is finished.

run.iterate_dataset_items would than be just thin wrapper, something like:

pseudo Python code:

def iterate_dataset_items(...):
  def is_finished():
    run = self.wait_for_finish(wait_duration=poll_interval, timeout=timeout)
    return run is None or run.status in _TERMINAL_STATUSES`
  
  yield from self.dataset().iterate_items(...,stop_condition_callback=is_finished)

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

t-tooling Issues with this label are in the ownership of the tooling team. tested Temporary label used only programatically for some analytics.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Live iteration over a run's dataset while the run is still running

3 participants