Repository navigation
fix(vector stores): refresh server poll intervals - #3775
Hughhhhcoder wants to merge 2 commits into
Conversation
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
sylvesterkaczmarek
left a comment
There was a problem hiding this comment.
I traced this through all four vector-store polling paths. The important part is that an omitted poll_interval_ms now remains omitted across iterations, so _get_poll_interval_ms(response.headers) is evaluated against each pending response instead of the first server hint becoming accidental caller state.
An explicit interval is still stable: it continues to set X-Stainless-Custom-Poll-Interval once and is selected on every iteration. When no server hint is present, _get_poll_interval_ms also preserves the existing 1000 ms fallback.
The two-response 100 ms -> 2000 ms regression directly pins the failure mode, and the shared parameterized test covers the sync/async file and file-batch entry points. I don't see a blocking behavior change beyond the intended response-scoped hint refresh.
|
Friendly follow-up at the one-week mark: this remains mergeable and ready for maintainer review. The parameterized regression covers all four sync/async file and file-batch polling paths, while preserving explicit intervals and the existing fallback. Please let me know if any adjustment or additional validation would help. |
sylvesterkaczmarek
left a comment
There was a problem hiding this comment.
Rechecked the unchanged head after the follow-up. The four vector-store polling paths preserve omitted intervals so each pending response can refresh the server hint, while explicit intervals and the 1000 ms fallback remain stable. I do not see a remaining blocker.
…ynamic-poll-hints # Conflicts: # src/openai/lib/_vector_stores.py
|
Resolved the conflict with current main after #3401 added bounded file-polling deadlines. The updated branch preserves both behaviors: omitted poll intervals are read from each pending response, while file polling still caps every sleep at the remaining deadline. Validation on the merged head: 72 focused polling/deadline tests passed in both the default and Pydantic v1 lanes; Ruff check, Ruff format check, and git diff --check also passed. This was a normal merge and push, with no history rewrite. |
sylvesterkaczmarek
left a comment
There was a problem hiding this comment.
Rechecked merge head 22f88524. The conflict resolution preserves both behaviors: omitted poll intervals are still refreshed from each pending response, and file polling still caps each sleep at the remaining deadline. The focused polling/deadline coverage is appropriate. No blocker from me.
|
Independent offline verification (AI-assisted): I extended the poll-hint check across all four actual helper paths (sync/async file and batch) with hint → missing hint → hint, hint → zero hint → hint, explicit fixed interval, and explicit zero interval. Two file-only cases check changing hints against a fake monotonic deadline: 18 cases total. Main Missing hints return to the one-second fallback even after a prior hint; server zero is read on its own response; explicit zero remains fixed. The deadline cases cap the second requested sleep at remaining time and raise before a further retrieve. Poll headers remain consistent and completed resources retain identity. This uses fake resource hooks and a fake clock: no actual sleeps, requests or latency measurements, and no claim about public-wrapper/provider behavior. Python 3.14.7, isolated source trees and shared dependency lane; network-denied. Focused evidence, not full SDK or merge approval. Reproducer: import asyncio,importlib,json
from types import SimpleNamespace
from openai._types import omit
source=importlib.import_module('openai.lib._vector_stores')
async def main():
rows=[];original=source.time.monotonic
try:
for batch,is_async in [(False,False),(False,True),(True,False),(True,True)]:
for name,hints,explicit,expected in [('hint_missing_hint',[100,None,2000],omit,[.1,1.,2.]),('hint_zero_hint',[100,0,200],omit,[.1,0.,.2]),('explicit_fixed',[100,None,2000],250,[.25,.25,.25]),('explicit_zero',[100,None,2000],0,[0.,0.,0.])]+([] if batch else [('deadline_dynamic',[100,2000],omit,[.1,1.])]):
clock={'now':0.};source.time.monotonic=lambda:clock['now'];sleeps=[];requests=[]
terminal=SimpleNamespace(status='completed',file_counts=SimpleNamespace(in_progress=0));pending=SimpleNamespace(status='in_progress',file_counts=SimpleNamespace(in_progress=1));responses=[]
for hint in hints:
headers={} if hint is None else {'openai-poll-after-ms':str(hint)};responses.append(SimpleNamespace(headers=headers,parse=lambda:pending))
responses.append(SimpleNamespace(headers={},parse=lambda:terminal))
def retrieve(*args,**kwargs):requests.append(dict(kwargs['extra_headers']));return responses.pop(0)
async def async_retrieve(*args,**kwargs):return retrieve(*args,**kwargs)
def sleep(seconds):sleeps.append(seconds);clock['now']+=seconds
async def async_sleep(seconds):sleep(seconds)
resource=SimpleNamespace(with_raw_response=SimpleNamespace(retrieve=async_retrieve if is_async else retrieve),_sleep=async_sleep if is_async else sleep)
function=getattr(source,('async_' if is_async else '')+'poll_vector_store_'+('file_batch' if batch else 'file'));kwargs={'vector_store_id':'fictional-store','poll_interval_ms':explicit}
if name=='deadline_dynamic':kwargs['max_wait_seconds']=1.1
error=None;result=None
try:
result=function(resource,'fictional-id',**kwargs)
if is_async:result=await result
except Exception as exc:error=type(exc).__name__
expected_error='TimeoutError' if name=='deadline_dynamic' else None
expected_headers={'X-Stainless-Poll-Helper':'true'}
if explicit is not omit:expected_headers['X-Stainless-Custom-Poll-Interval']=str(explicit)
ok=error==expected_error and [round(x,9) for x in sleeps]==expected and all(r==expected_headers for r in requests) and (expected_error is not None or result is terminal)
rows.append({'case':name,'batch':batch,'async':is_async,'sleeps':sleeps,'expected':expected,'error':error,'retrieve_calls':len(requests),'pass':ok})
finally:source.time.monotonic=original
print(json.dumps({'source_module':source.__file__,'cases':rows,'passed':sum(r['pass'] for r in rows),'failed':sum(not r['pass'] for r in rows),'scope':'Actual polling helpers with fake resource hooks and fake monotonic clock; no actual sleep, HTTP request, deadline timing or public-wrapper proof.'}))
asyncio.run(main()) |
Changes being requested
When
poll_interval_msis omitted, the vector-store polling helpers currently assign the first response'sopenai-poll-after-msvalue back topoll_interval_ms. That makes the value appear caller-specified on later iterations, so subsequent server hints are ignored.For example, two pending responses with 100 ms and 2000 ms hints currently sleep for
[0.1, 0.1]seconds instead of[0.1, 2.0].This change keeps an explicitly supplied interval fixed, while resolving an omitted interval from every pending response. The behavior now matches the response-scoped hint handling in the other OpenAI SDK polling implementations.
The regression covers all four public vector-store polling paths:
FilesAsyncFilesFileBatchesAsyncFileBatchesAdditional context & links
The patch is limited to the SDK-owned
src/openai/lib/_vector_stores.pyhelper and its focused tests. It does not change any public signature, request header, terminal-state behavior, or generated resource file.Validation:
[0.1, 0.1]../scripts/lint: Ruff, Pyright, mypy, and import check passedgit diff --check: passed