-
Notifications
You must be signed in to change notification settings - Fork 29.3k
Pull requests: apache/spark
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
[SPARK-58559][PYTHON] Package all of sbin in PySpark classic distribution
#57763
opened Aug 4, 2026 by
nchammas
Contributor
Loading…
[SPARK-58558][SQL] Remove requireAllClusterKeysForCoPartition as a gate for SPJ
#57762
opened Aug 4, 2026 by
pan3793
Member
Loading…
[WIP][ML] Avoid duplicate StringIndexer skip lookups
#57761
opened Aug 4, 2026 by
zhengruifeng
Contributor
•
Draft
[SPARK-58207][SQL][FOLLOWUP] Skip runtime filter pushdown for nondeterministic filters
#57760
opened Aug 4, 2026 by
peter-toth
Contributor
Loading…
[SPARK-49828][SQL] Make Column(expression) usable outside of the org.apache.spark package
#57759
opened Aug 4, 2026 by
peter-toth
Contributor
Loading…
[SPARK-58553][PS] Use native Spark functions for NumPy fmax and fmin
#57758
opened Aug 4, 2026 by
zhengruifeng
Contributor
Loading…
[SPARK-58557][ML] Reuse ML vector data type singleton
#57757
opened Aug 4, 2026 by
zhengruifeng
Contributor
Loading…
[WIP][MLLIB] Rewrite CountVectorizer fitting with DataFrame APIs
#57755
opened Aug 4, 2026 by
zhengruifeng
Contributor
•
Draft
[SPARK-58556][CONNECT][UI] Show ML cache status in Spark Connect UI
#57754
opened Aug 4, 2026 by
zhengruifeng
Contributor
•
Draft
[SPARK-58549][SQL] Preserve key-grouped partitioning and ordering across a DSv2 scan merge
#57753
opened Aug 4, 2026 by
peter-toth
Contributor
Loading…
[SPARK-58551][PYTHON] Python Data Sources Limit Pushdown API
#57752
opened Aug 4, 2026 by
ganeshashree
Contributor
Loading…
[SPARK-58552][SQL][UI] Add total task time column to the SQL / DataFrame tab
#57751
opened Aug 4, 2026 by
ulysses-you
Contributor
Loading…
[SPARK-58550][ML] Delay GaussianMixture aggregation allocations
#57750
opened Aug 4, 2026 by
zhengruifeng
Contributor
•
Draft
[SPARK-58548][PS] Use native Spark function for NumPy heaviside
#57749
opened Aug 4, 2026 by
zhengruifeng
Contributor
Loading…
[SPARK-36284][CORE][SHUFFLE] Add shuffle checksum support for push-based shuffle
#57748
opened Aug 4, 2026 by
Dreamstick9
Loading…
[SPARK-58547][CONNECT] Expose operation IDs for end-to-end request attribution
#57747
opened Aug 4, 2026 by
cloud-fan
Contributor
Loading…
[SPARK-58544][SQL] Fix vector distance and norm functions returning wrong results from intermediate float overflow
#57746
opened Aug 4, 2026 by
SEPURI-SAI-KRISHNA
Loading…
[SPARK-58511][SQL] Bypass ineffective pre-shuffle partial aggregation at runtime
#57742
opened Aug 4, 2026 by
ulysses-you
Contributor
Loading…
[SPARK-58538][INFRA] Add branch-4.3 CI scheduler and release integration
#57741
opened Aug 4, 2026 by
HeartSaVioR
Contributor
Loading…
[SPARK-58536] Fix getDefaultFinalStatus to return SUCCEEDED in cluster mode
#57736
opened Aug 4, 2026 by
Lobo2008
Loading…
[SPARK-56573][SQL] Use a full-range non-negative random seed for unseeded sampling
#57732
opened Aug 4, 2026 by
stanyao
Contributor
Loading…
[SPARK-58531][SQL][CONNECT] Make spark.sql.artifact.copyFromLocalToFs.allowDestLocal a static conf
#57730
opened Aug 3, 2026 by
haoyangeng-db
Contributor
Loading…
[SPARK-58015][INFRA][FOLLOWUP] Make dependency group names compliant with the PyPA specification
#57729
opened Aug 3, 2026 by
ueshin
Member
Loading…
[SPARK-58529][PYTHON] Unify RESULT_ROWS_MISMATCH message and consolidate row count verification in worker.py
#57728
opened Aug 3, 2026 by
Yicong-Huang
Contributor
Loading…
[SPARK-58523][SQL] Add Catalyst runtime filtering interface for DSv2 scans
#57727
opened Aug 3, 2026 by
szehon-ho
Member
Loading…
Previous Next
ProTip!
Follow long discussions with comments:>50.