← All tasks
pythonopenai/openai-agents-python #4472Not a task: already works

tests-windows pins one Python version, so clock-resolution bugs are invisible on Windows for 3 of 5 supported versions

envgap__openai__openai-agents-python-4472

01 / FAILURE SIGNATURE

As reported upstream

No identifying execution failure has been captured.
Not a benchmark task.
  • The project already builds and runs before the fix, so there is nothing to repair.

02 / ENVIRONMENT RECIPE

Base commit
37a7aa20cee5f16d3720214c39dc66ca9f143e74
Manifest
pyproject.toml
Reproduce
Awaiting issue-specific recipe
Run under trace
Awaiting a meaningful runtime command

03 / ORIGINAL ISSUE TEXT

openai/openai-agents-python #4472 · read the original issue
`tests-windows` pins a single Python version, so a whole class of bug is invisible to CI on the one platform where it occurs. #4392 was an instance of that class; this is about the gap that let it sit there.

## The gap

```
tests (matrix)   ubuntu-latest    3.10  3.11  3.12  3.13  3.14
tests-windows    windows-latest   3.13 only
packaged-contract-windows         3.13 only
```

The project's classifiers claim 3.10 through 3.14, and the Linux matrix honours that. On Windows only one version is exercised, and it happens to be the version where the failure in #4392 does not reproduce.

## Why that particular version hides things

From the measurements in #4392, `datetime.now()` resolution on Windows over a 1000-call burst:

| Python | distinct timestamps per 1000 calls |
|---|---:|
| 3.10.19 | 2 |
| 3.11.14 | 2 |
| 3.12.12 | 2 |
| 3.13.11 | 1000 |

3.13 is the version where the Windows clock stopped being coarse. Pinning Windows CI to exactly that version means anything sensitive to clock granularity, ordering on tied timestamps, or short-interval timing is untested on Windows for three of the five supported versions. In #4392 that was 37 tests failing on a clean checkout while CI stayed green.

To be clear about scope: #4392 fixed the instance, so nothing is failing right now as far as I know. This is about the next one.

## Suggestion, sized for the CI bill

Windows runners bill at a higher multiplier, so a full five-version Windows matrix is probably not worth it. Adding one older version would cover the class:

```yaml
tests-windows:
  runs-on: windows-latest
  strategy:
    matrix:
      python-version: ["3.10", "3.13"]
```

3.10 is the declared minimum and sits on the coarse-clock side; 3.13 is what runs today. That is 2x the current Windows cost rather than 5x, and it would have caught #4392 before it reached a contributor's machine.

If even that is too much, running the older version on a schedule rather than per-PR would still surface the class, just later.

I am happy to send the PR if you tell me which shape you want, or to leave it if the cost is not worth the coverage. Filing this rather than opening a PR straight away because the CI budget is your call, not mine.
Continue on GitHub ↗

04 / LABELS

Labels from the report text only; not yet run

No supported category has been assigned.

Label rules and the text that matched
[]