Kubernetes v1.37.0 is live!

29 views
Skip to first unread message

Agustina Barbetta

unread,
Aug 26, 2026, 12:36:28 PM (7 days ago) Aug 26
to kubernetes-announce, dev
Kubernetes Community,

Kubernetes v1.37.0 has been built and pushed using Golang version 1.26.5.

The release notes have been updated in CHANGELOG-1.37.md, with a pointer to them on GitHub:


v1.37.0

Documentation

Downloads for v1.37.0

Source Code

filename sha512 hash
kubernetes.tar.gz 840eab917c60a35306cfbfa40b3b6a70c10723f4326e5a0aa45084a83d3e45ca5def613e7fbd0dd53955ef091ed672c3035f5f42be5540099992f8723874a80b
kubernetes-src.tar.gz 639781033f79b2a795f0557cbd19013e630ab26b00aad6762d9a94100e87ac4dd5aaa3a08513948ed564bf71f8fb170404310612b3454c9ed1021c8333ed3525

Client Binaries

filename sha512 hash
kubernetes-client-darwin-amd64.tar.gz 6f9a612f4bd379b049b744a64e94eefc7a952fdf9b4097e09c74f8a0c9e7d3c6bd8645aea678079d4208d34ec33f200c10ce90812f08bae3c439f982dd5f0015
kubernetes-client-darwin-arm64.tar.gz 199721e27425b1de3e95f260da1e61d76dc7185ad7cba55e219b9e9a6b133fe6a92590b81d784752f08e3b7d03676d024a92b742e48628d0f6a87f9fec7f8047
kubernetes-client-linux-386.tar.gz a42ba1f363d1afb89f61ee09250faf60a754cbea733aa346296b5a35504c6375d656162d16bd534cb442e2b93a314c3b77de82121268f5bff30aec3f8ae8125c
kubernetes-client-linux-amd64.tar.gz ff61f94cf73281e24b8b7f16bd7ae15a928ef886399a2e8a360c170828babade73aeb284d37c596fc6c7dc68876b1550b714c1fc3ac4263adb8aef9e30a7c42f
kubernetes-client-linux-arm.tar.gz a93e9c857ac63d30fb08a59afc22630f65c2ab65db0b840056b1c2fb7d8c54c0e42c862e2215f4bbba03d954f90b66fe1323d6a9b77585499f54838e90bd4b84
kubernetes-client-linux-arm64.tar.gz 28167564b99d2440c3be873aff2f6bb2726ba55b58a3af3994f877f3d680e539d632e5162f96838d9efd06b87f5e1dbc811782f01f86a221afd90bf7fa9d0328
kubernetes-client-linux-ppc64le.tar.gz c40251ca9c0944c748e510feabce024a5e822c5868bb4922223415c7979589482702db22a0fd8976bb19dd5652bec41ce69c037daf2fe2cb3290ca8733149a4b
kubernetes-client-linux-s390x.tar.gz 9bdab3d71a426a4b5371f7a8e99621ec9f6c92bb83700006b705b3a32ea33d6792fc5506fd08fc09479b91f1fa195c3d1c74765957e41b67a906b7f4b1317453
kubernetes-client-windows-386.tar.gz 483b14fe5a980c975d8c3dcd558eaf72937f827b3a844792baff30cf2201600eacf7fa8b6c1986d6699b8bcccd1e6c9aa057f2eb54c2b87444055d786927919b
kubernetes-client-windows-amd64.tar.gz e9368652abca1590083b67be9d8f763aa27a4025d2a30a56d779ebb3150f062dc21d6673ed6ebbe04afe2a4ab4345276de67dac8e01081e9d0d199b7a492801f
kubernetes-client-windows-arm64.tar.gz 7e31107af1763d0acf2eaea083e4a64c42d524205827ea3f7005043c9aeed61b2197d5c0021d3679f05530f03907c8a241eea4dcde19038492609bda0ebfc044

Server Binaries

filename sha512 hash
kubernetes-server-linux-amd64.tar.gz f537143ccadab75bce7cfa97ae29999ec4fa5e09405ad65cb9b1dcd3b132a0eefde4fcc8653e0059113b1ffb163255df26cd4a725b07601a261c94e17cbcfc6e
kubernetes-server-linux-arm64.tar.gz 09d9be33f71a93704da6684e0c88872d3140e55c4ad9ae87a92045ac0423ce3770b282b6d5f13fbdb5e49dd817385d0e261ba5f22637eb14712f5301d39f8795
kubernetes-server-linux-ppc64le.tar.gz 56e56d0bfa217a40987d0cb6a181ffaed35bf7efcb13856fe681dc507728b3d71f39c42e021e20732c0f751cedfb61b00077f388f26273d4c4007a5e0952dbcd
kubernetes-server-linux-s390x.tar.gz a31dc4c535a2a064581604c82ce7086c79fe4826b1cdf276ce714b9607d682514f3ed7c1059b6f1feae9111ebdef534b34b571cce3b5c165447385074ba6f659

Node Binaries

filename sha512 hash
kubernetes-node-linux-amd64.tar.gz f2b47175d2a2844ae1673d239b7b8be18b7dc2363077925bba654802aa9f35aedafb7086aae9005545267aef21518986177fd7ca4b25edfe6e400102adfccad2
kubernetes-node-linux-arm64.tar.gz 3c7d592a9da9466c41ae0f913215bbbe7e62704e83b1c5c38c0bbcf9ae20de4c62186e4ab638bcb2cd0d4c1a658d945dbb634413409fb601da0231f47bed219c
kubernetes-node-linux-ppc64le.tar.gz 7c0a1d8aed394884e15b6e50701237fdd6664aa6fe302f34bce3341918cb3acc29df38084794d6f4d08b42bde073da84371550b2b444d40e746a3be63653ceda
kubernetes-node-linux-s390x.tar.gz 71d4ee0c4cfac965cfa9054bd3ccc1030dfbc454af6e56c8ca63e83d6327fff10e5561d39c69cc16ef875cd499f30673fbec3aa770d51ef6d1f4414b1040b9b9
kubernetes-node-windows-amd64.tar.gz 50ce08d0cac9fe0c772cc87a8097f0bdc1866ddbf25cc2842d91e929f079b984836b0a2bb427b0b34e33eec44502afacbfde74d06acbc421ebe85287ae7df1b4

Container Images

All container images are available as manifest lists and support the described architectures. It is also possible to pull a specific architecture directly by adding the "-$ARCH" suffix to the container image name.

name architectures
registry.k8s.io/conformance:v1.37.0 amd64, arm64, ppc64le, s390x
registry.k8s.io/kube-apiserver:v1.37.0 amd64, arm64, ppc64le, s390x
registry.k8s.io/kube-controller-manager:v1.37.0 amd64, arm64, ppc64le, s390x
registry.k8s.io/kube-proxy:v1.37.0 amd64, arm64, ppc64le, s390x
registry.k8s.io/kube-scheduler:v1.37.0 amd64, arm64, ppc64le, s390x
registry.k8s.io/kubectl:v1.37.0 amd64, arm64, ppc64le, s390x

Changelog since v1.36.0

Urgent Upgrade Notes

(No, really, you MUST read this before you upgrade)

  • ACTION REQUIRED: Graduated the SELinuxMount feature gate to GA. The feature is enabled by default in v1.37, which may break existing workloads in clusters with SELinux enabled. Please see the Kubernetes blog post to identify potentially problematic workloads in a v1.36 cluster and how to fix them or opt out of SELinuxMount changes before upgrading to v1.37. Admins of clusters without SELinux enabled can ignore this release note. (#139956, @jsafrane) [SIG API Machinery, Apps and Node]
  • ACTION REQUIRED: Converted the DisruptionMode enum field to a struct to support future extensibility. Promoted the scheduling.k8s.io API group from v1alpha2 to v1alpha3 and dropped v1alpha2 entirely. Remove all v1alpha2 objects from the kube-apiserver before performing the cluster update. (#138572, @dom4ha) [SIG API Machinery, Apps, CLI, Etcd, Node, Scheduling and Testing]
  • ACTION REQUIRED: Fixed eventRecordQPS handling in kubelet configuration to treat 0 as unlimited (no rate limit), aligning behavior with the documentation. Users relying on the previous default behavior should explicitly set a non-zero value (for example, 50). (#117119, @HirazawaUi) [SIG API Machinery, Auth and Node]
  • ACTION REQUIRED: Updated the kubelet to log its effective configuration at startup. Because these logs can expose configuration details, cluster administrators should restrict the nodes/logs ClusterRole to trusted users. This is largely a reminder of an existing best practice, since most effective configuration values can already be inferred from other log messages or kubelet behavior. (#139837, @SergeyKanzhelev) [SIG Node]

Changes by Kind

Dependency

  • Updated google.golang.org/grpc to v1.82.1, which adds a server-side limit on HTTP/2 control frame flooding. It also removes the GRPC_GO_EXPERIMENTAL_DISABLE_STRICT_PATH_CHECKING environment variable, so strict path checking is always on. (#140740, @dims) [SIG API Machinery, Architecture, Auth, CLI, Cloud Provider, Network, Node and Scheduling]
  • Updated the kubelet's embedded cAdvisor to use the leaner github.com/google/cadvisor/lib module, which removes three long-deprecated or legacy surfaces:
    • Deprecated cAdvisor flags are no longer accepted and the kubelet will fail to start if any are set (only --housekeeping-interval is kept): --application-metrics-count-limit, --boot-id-file, --container-hints, --containerd, --containerd-namespace, --enable-load-reader, --event-storage-age-limit, --event-storage-event-limit, --global-housekeeping-interval, --log-cadvisor-usage, --machine-id-file, --storage-driver-user, --storage-driver-password, --storage-driver-host, --storage-driver-db, --storage-driver-table, --storage-driver-secure, --storage-driver-buffer-duration. Remove these from your kubelet configuration.
    • cAdvisor application/custom metrics are no longer collected: the userDefinedMetrics field in /stats/summary and the custom container_application_* families in /metrics/cadvisor.
    • The /metrics/cadvisor series container_cpu_load_average_10s, container_cpu_load_d_average_10s, and container_tasks_state are no longer exported. (#139870, @dims) [SIG API Machinery, Auth, Instrumentation, Network, Node and Testing]
  • Updated the default etcd version to v3.7.0-rc.0. (#139427, @Jefftree) [SIG API Machinery, Cloud Provider, Cluster Lifecycle, Etcd and Testing]
  • Updated the default etcd version to v3.7.0. (#140333, @Jefftree) [SIG API Machinery, Auth, Cloud Provider, Cluster Lifecycle, Etcd, Node, Scheduling and Testing]
  • Updated the etcd client library to v3.6.10. (#138393, @humblec) [SIG API Machinery, Architecture, Auth, CLI, Cloud Provider, Cluster Lifecycle, Etcd, Instrumentation, Network, Node, Scheduling and Storage]

Deprecation

  • Changed kubeadm to explicitly set KubeProxyConfiguration.mode to iptables when KubeProxyConfiguration is not provided or when the mode field is empty. In v1.37, kube-proxy will warn if the field is not explicitly set as part of the planned transition to nftables as the default mode in a future release. (#139777, @neolit123) [SIG Cluster Lifecycle]
  • Deprecated the ignored --filename/-f flag on kubectl run. (#138671, @Suknna) [SIG CLI]
  • Locked the deprecated DeclarativeValidationTakeover feature gate to its default value; it can no longer be set. (#139212, @yongruilin) [SIG API Machinery]
  • kubeadm: Added a (delayed) warning that kube-proxy's ipvs mode is deprecated since v1.35 and users on newer Linux kernels should be using the nftables mode instead, which became GA in v1.33. For older kernel versions, users can use iptables, which is still the default. (#139067, @neolit123) [SIG Cluster Lifecycle]

API Change

  • Added Alpha support for DRA device compatibility groups, guarded by the DRADeviceCompatibilityGroups feature gate (disabled by default). DRA drivers can declare opaque compatibilityGroups on each device.consumesCounters[] entry of a ResourceSlice, and the scheduler only co-allocates devices drawing from the same counter set when their declared groups intersect. This moves detection of incompatible co-allocation from preparation-time failure to scheduling-time rejection. (#139795, @omeryahud) [SIG API Machinery, Node, Scheduling and Testing]
  • Added Alpha support for binding service account tokens to webhook configurations with attestations, behind the APIServerWebhookAuthenticationToken feature gate. This enables API servers to authenticate to admission webhooks with scoped tokens. (#140113, @pmengelbert) [SIG API Machinery, Apps, Auth and Testing]
  • Added Alpha support for defining the file owner of atomically written volume files, behind the AtomicWriteVolumeUserFields feature gate (disabled by default). (#139764, @gavinkflam) [SIG API Machinery, Apps, Auth, Storage and Testing]
  • Added CompositePodGroup support to the building block APIs and the workloadbuilder library. (#140717, @helayoty) [SIG API Machinery, Apps, Auth, Scheduling and Testing]
  • Added Workload-aware scheduling (WAS) support to the Job controller by integrating it with the workloadbuilder library and the Workload building block APIs. (#140188, @helayoty) [SIG API Machinery, Apps, Auth, Network, Node, Scheduling, Storage and Testing]
  • Added CheckpointPod and RestorePod RPCs to the CRI v1 RuntimeService API for Pod-level checkpoint and restore. (#140366, @rst0git) [SIG Node, Testing and Windows]
  • Added a PreemptionPolicy field to PodGroup, allowing users to control how kube-scheduler preempts Pods that belong to a PodGroup during workload-aware preemption. (#139240, @ania-borowiec) [SIG Scheduling]
  • Added a protocol field to httpGet probes so that liveness, readiness, and startup probes can run over HTTP/2 cleartext (H2C). (#139429, @amritansh1502) [SIG API Machinery, Apps, Node and Testing]
  • Added a defense-in-depth check to the NodeRestriction admission plugin for PodCertificateRequests. A node can only create a PodCertificateRequest referring to a particular signer name if the Pod actually mounts a podCertificate projected volume source that refers to that signer name, or if an authorization check for the user (with verb=request-podcertificate-signer and resource=<signerName>) succeeds. (#140006, @ahmedtd) [SIG Auth and Testing]
  • Added a second Alpha of DRA resource availability visibility (KEP-5677), behind the DRAResourcePoolStatus feature gate (disabled by default). The ResourcePoolStatusRequest controller counts partitionable and consumable devices correctly, counting each device once, ignoring AdminAccess, and treating taints as unavailable. Optional fields describe partition and shareable availability. These accounting fixes change the numbers reported in v1.36. (#140170, @nmn3m) [SIG API Machinery, Apps, Auth, Node and Testing]
  • Added an opt-in userspace TCP proxy to the nftables kube-proxy backend to serve localhost NodePort Services on IPv4 and IPv6. (#138427, @AustinAbro321) [SIG Instrumentation, Network and Testing]
  • Added dry-run support to unsafe corrupt object deletion, letting administrators test deletion operations safely before running them. (#134037, @ibihim) [SIG API Machinery and Testing]
  • Added generated declarative validation functions for Go consumers of the DRA device metadata v1alpha1 API. (#140687, @alaypatel07) [SIG Node]
  • Added scheduler support for preempting lower-priority Pods to make room for Deferred in-place Pod resizes of higher-priority Pods when the InPlacePodVerticalScalingSchedulerPreemption Alpha feature gate is enabled. (#140000, @natasha41575) [SIG API Machinery, Apps, Node, Scheduling, Storage and Testing]
  • Added support for derived attributes in DRA, allowing claims to define virtual attributes using CEL expressions and use them in device constraints. This enables co-allocation of devices across different domains, for example GPUs and NICs on the same NUMA node, even if their drivers publish physical attributes differently. (#140029, @gauravkghildiyal) [SIG API Machinery, Node, Scheduling and Testing]
  • Added support for dynamically resizing memory-backed volumes behind the Alpha InPlacePodVerticalScalingMemoryBackedVolumes feature gate. (#139425, @natasha41575) [SIG API Machinery, Apps, Autoscaling, CLI, Node, Scheduling, Storage and Testing]
  • Added support for selecting ResourceSlices by pool name with the field selector spec.pool.name. (#138456, @yaroslavborbat) [SIG API Machinery, Node and Testing]
  • Added support for setting Unix permission bits (0000-01777) through the mode field on emptyDir volume directories at creation time. (#140244, @nispriha) [SIG API Machinery, Apps, Node, Storage and Testing]
  • Added support for specifying bind mount options (noexec, nodev, nosuid) per container volume mount. (#140013, @nispriha) [SIG API Machinery, Apps, Autoscaling, Node, Scheduling and Testing]
  • Added the API changes required for reporting volume health. (#140194, @gnufied) [SIG API Machinery, Apps, Architecture, Auth, Etcd, Instrumentation, Node, Storage and Testing]
  • Added the CompositePodGroup API to scheduling.k8s.io/v1alpha3. (#139596, @jdzikowski) [SIG API Machinery, Apps, Auth, Etcd, Node, Scheduling and Testing]
  • Added the Recreate update strategy for StatefulSet, mirroring the Deployment Recreate strategy by deleting all Pods, waiting for them to terminate completely, and then creating Pods according to the podManagementPolicy field. (#137187, @galal-hussein) [SIG Apps and Testing]
  • Added the --concurrent-disruption-syncs flag to kube-controller-manager to configure the number of concurrent disruption controller workers. (#140014, @xigang) [SIG API Machinery, Apps, Auth and Testing]
  • Added the .spec.evictionResponders Pod field, along with the EvictionRequest and Eviction resources. A set of requesters and responders can use these to coordinate graceful eviction of Pods. (#137050, @atiratree) [SIG API Machinery, Apps, Architecture, Auth, CLI, Etcd and Testing]
  • Added the DefaultPodSysctls kubelet configuration field for default Pod sysctls on Linux nodes, behind the Alpha DefaultPodSysctls feature gate (disabled by default). (#140052, @VeraQin) [SIG Node and Testing]
  • Added the DisruptionMode and PreemptionPolicy fields to the Workload and CompositePodGroup APIs to support workload-aware preemption for CompositePodGroups. (#140634, @tosi3k) [SIG API Machinery, Auth, Etcd, Node, Scheduling and Testing]
  • Added the GracefulNodeShutdownInProgress, DrainInProgress, Drained, MaintenancePlanned, and MaintenanceInProgress Node lifecycle conditions. (#139993, @rthallisey) [SIG Apps and Node]
  • Added the PodGroupPostFilter extension point to the scheduling framework, replacing the internal hardcoding for WorkloadAwarePreemption with a configurable extension point for PodGroups. (#139674, @GFilipek) [SIG Scheduling and Testing]
  • Added the PreemptionPolicy field to PodGroupTemplate to define the policy for workload-aware preemption. (#140312, @ania-borowiec) [SIG API Machinery, Apps, Scheduling and Testing]
  • Added the SchedulerPreQueueingHints feature gate (Alpha, disabled by default). When enabled, scheduler plugins can provide a PreQueueingHintFn that narrows the set of Pods evaluated on cluster events, improving scheduling throughput. The DRA plugin implements this to optimize ResourceClaimTemplate-based workloads. (#138916, @geetasg) [SIG API Machinery, Apps, Architecture, Auth, CLI, Cloud Provider, Instrumentation, Network, Node, Scheduling, Storage, Testing and Windows]
  • Added the core machinery for Conditional Authorization, which enables authorizers to allow requests conditionally based on the request's content, rather than only allowing or denying them outright. (#137513, @luxas) [SIG API Machinery, Auth, Node and Testing]
  • Allowed HorizontalPodAutoscaler conditions to optionally include the observedGeneration at the time the condition was recorded. (#138653, @adrianmoisey) [SIG API Machinery, Apps, Autoscaling and Testing]
  • Changed the kube-apiserver CBOR encoder to encode collections item by item instead of all at once. (#138808, @chenk008) [SIG API Machinery, Apps, Auth, Autoscaling, CLI, Cloud Provider, Cluster Lifecycle, Contributor Experience, Instrumentation, Network, Node, Release, Scalability, Scheduling, Storage, Testing and Windows]
  • DRA: Added Alpha support for DRAOptionalNodeOperations (the SkipNodeOperations field in ResourceSlice and ResourceClaim), which allows skipping node-level preparation and cleanup operations. (#139933, @troychiu) [SIG API Machinery, Autoscaling, Instrumentation, Node, Release, Scheduling and Testing]
  • Fixed CEL cost estimation for metadata.name and metadata.generateName in CRDs to correctly account for the default maximum length of 253 characters, unless additional validations are defined on those metadata fields. (#139573, @jpbetz) [SIG API Machinery]
  • Fixed DRA CapacityRequestPolicyRange to support fractional quantities in milli-scale. (#140161, @sunya-ch) [SIG API Machinery, Node and Scheduling]
  • Fixed Pod status validation for reported Linux container user UIDs to accept values above 2147483647 and up to the unsigned 32-bit UID limit. (#138574, @Kunalbehbud) [SIG Apps and Node]
  • Fixed a v1.34+ regression handling containers with environment values set from Secret API objects containing binary non-utf8 data. (#139168, @liggitt) [SIG Architecture, Node and Testing]
  • Fixed a bug in DRA consumable capacity where the DistinctAttribute constraint was not correctly enforced for each device when a request allocated multiple devices. (#140600, @GunaKKIBM) [SIG API Machinery, Apps, CLI, Etcd, Network, Node, Release, Scheduling and Testing]
  • Fixed the overestimation of a Pod's resource footprint during resize operations for multi-container Pods. (#140047, @natasha41575) [SIG Node and Scheduling]
  • Graduated Pod Certificates to GA, with the PodCertificateRequest feature gate enabled by default. The PKIXPublicKey and ProofOfPossession fields, deprecated in PodCertificateRequest v1beta1, are removed from the v1 API. Update clients that still set these fields before upgrading. (#139579, @yt2985) [SIG API Machinery, Apps, Architecture, Auth, Etcd, Node, Scheduling and Testing]
  • Graduated Pod hostname overrides to GA. The HostnameOverride feature gate is locked to enabled. (#139116, @HirazawaUi) [SIG API Machinery, Apps, Node and Testing]
  • Graduated the ClusterTrustBundle and ClusterTrustBundleProjection feature gates to GA and enabled them by default. (#139437, @stlaz) [SIG API Machinery, Apps, Architecture, Auth, Etcd, Node, Storage and Testing]
  • Improved CEL error messages in Dynamic Resource Allocation to provide guidance when accessing non-existent device attributes. Error messages link to documentation on handling optional fields using orValue() and has(). (#136709, @gzb1128) [SIG API Machinery, Node and Scheduling]
  • Moved the NodeSyncPeriod field from KubeCloudSharedConfiguration to CloudControllerManagerConfiguration.NodeLifecycleController.NodeMonitorPeriod. (#137964, @niewysoki) [SIG API Machinery, Apps, Cloud Provider, Instrumentation and Node]
  • Promoted DRA Workload resource claims to Beta. The DRAWorkloadResourceClaims feature gate remains disabled by default. (#140334, @nojnhuh) [SIG API Machinery, Apps, Etcd, Node, Scheduling and Testing]
  • Promoted kubelet volume metrics (storage_operation_duration_seconds, volume_operation_total_seconds) from Alpha to Beta stability, providing stronger API and label stability guarantees for metric consumers. (#136189, @bhope) [SIG Instrumentation and Storage]
  • Promoted the DRA Device Taints and Tolerations feature to GA, making it available via the resource.k8s.io/v1 API. (#138676, @pohly) [SIG API Machinery, Apps, Architecture, Auth, Etcd, Node, Scheduling, Storage and Testing]
  • Promoted the DRA extended resource feature to GA in v1.37. (#138488, @yliaog) [SIG API Machinery, Apps, Node, Scheduling and Testing]
  • Promoted the DRA metadata API to Beta. DRA driver authors that enable the feature must explicitly select which versions to support in their metadata output. (#140722, @pohly) [SIG Node and Testing]
  • Promoted the DRAResourceHealth kubelet gRPC API to v1; the schema is unchanged from v1alpha1. DRAPlugin.WatchHealthStatus is a mandatory method on the DRAPlugin interface in the k8s.io/dynamic-resource-allocation/kubeletplugin helper, replacing the optional versioned gRPC interface. This is a one-time Go API break: existing drivers must add the method to compile. Drivers without health support return ErrHealthNotSupported from it, or disable the service with HealthService(false). The helper serves both v1 and v1alpha1 by default, so drivers report health on kubelet v1.36 and older without extra configuration. The kubelet prefers v1 and, for three releases of transition, still consumes v1alpha1 from drivers that shipped before v1 existed. The v1alpha1 DRAResourceHealth API is deprecated and is planned to be removed in v1.40. The kubelet only opens the device health stream for plugins that advertise the service. (#139477, @harche) [SIG Node and Testing]
  • Promoted the HPAConfigurableTolerance feature gate to GA. (#140107, @jm-franc) [SIG API Machinery, Apps, Autoscaling and Testing]
  • Promoted the KubeletInUserNamespace feature gate to Beta. (#134639, @AkihiroSuda) [SIG API Machinery, Apps, Node and Testing]
  • Promoted the MemoryQoS feature gate to Beta. memoryThrottlingFactor defaults to nil, and memory.high is not set unless explicitly configured. (#140007, @QiWang19) [SIG Node and Testing]
  • Promoted the NodeDeclaredFeatures feature gate to GA. (#139763, @pravk03) [SIG Apps, Autoscaling, Node, Scheduling and Testing]
  • Promoted the PersistentVolumeClaimUnusedSinceTime feature gate to Beta in v1.37 and enabled it by default. PersistentVolumeClaims report the Unused condition, indicating how long a PersistentVolumeClaim has been unused to help identify candidates for cleanup. (#139620, @RomanBednar) [SIG Apps]
  • Promoted the StorageVersionMigration feature gate to GA. The storagemigration.k8s.io/v1 API group is enabled by default. (#138560, @michaelasp) [SIG API Machinery, Apps, Architecture, Auth, Etcd and Testing]
  • Promoted the VolumeLimitScaling feature gate, which adds the preventPodSchedulingIfMissing field to CSIDriver to prevent Pod scheduling on nodes missing a required CSI driver, to Beta. (#140612, @gnufied) [SIG API Machinery, Storage and Testing]
  • Promoted the metrics.k8s.io API from v1beta1 to v1 without changes. (#139223, @tico88612) [SIG Instrumentation]
  • Promoted the core Workload-Aware Scheduling (WAS) API types Workload and PodGroup to scheduling.k8s.io/v1beta1. Remove all v1alpha2 objects from the kube-apiserver before upgrading from v1.36 to v1.37. (#140184, @tosi3k) [SIG API Machinery, Apps, Auth, Etcd, Node, Scheduling and Testing]
  • Relaxed container security context validation so that updates to Pods may set allowPrivilegeEscalation together with CAP_SYSADMIN. Pod creation still rejects this combination. (#138834, @haircommander) [SIG Apps]
  • Removed the GangScheduling and WorkloadAwarePreemption feature gates. Use the GenericWorkload feature gate to enable core workload-aware scheduling functionality. (#139520, @macsko) [SIG API Machinery, Node, Scheduling and Testing]
  • Removed the generally available feature gate AnyVolumeDataSource, which was locked and enabled since v1.33. (#135336, @carlory) [SIG API Machinery, Apps, Storage and Testing]
  • Removed the unused PodStatusResult type from the Kubernetes API. This type had no REST endpoint and has been unused since 2015. (#136271, @adityasharmawork) [SIG API Machinery, Apps, Node and Testing]
  • Renamed signal enum keys in cri-api, which are prefixed with SIGNAL_ in the api.proto definition, to avoid conflicts with C++ macros. The wire format is unchanged. (#139251, @SergeyKanzhelev) [SIG Apps, Node and Testing]
  • Renamed the PodGroup condition PodGroupScheduled to PodGroupInitiallyScheduled to clarify that the condition is set only after a PodGroup is first scheduled successfully and may not reflect the PodGroup's current scheduling state. (#139743, @antekjb) [SIG API Machinery, Scheduling and Testing]
  • Switched the json tag for inlined TypeMeta fields in the API Go types from ",inline" to "". inline was not a recognized json serializer option and did not modify marshal or unmarshal behavior. (#138260, @liggitt) [SIG API Machinery, Apps, Architecture, Auth, CLI, Cloud Provider, Cluster Lifecycle, Etcd, Instrumentation, Network, Node, Scheduling, Storage and Testing]
  • Updated CDI spec version selection to be dynamic, preventing the generation of incompatible CDI specifications. (#137699, @alaypatel07) [SIG Apps, Node, Scheduling and Testing]
  • Updated Pod-level resource handling so that Pod resources determine QoS only when they include a resource request or limit. Empty Pod resources ({}, {requests:{}}, or {limits:{}}) no longer affect QoS calculation. (#137150, @KevinTMtz) [SIG Apps, CLI, Node and Scheduling]
  • Updated PodGroup and PodGroupTemplate to allow modifying minCount after creation. Changes to a PodGroupTemplate do not affect existing PodGroups. (#139279, @antekjb) [SIG API Machinery, Scheduling and Testing]
  • Updated the Alpha DRANodeAllocatableResources feature:
    • ResourceSlice mappings and PodStatus support direct allocations, for DRA drivers modeling CPU, memory, and hugepages as a resource, and overhead allocations, for accelerator host overhead.
    • The kubelet accounts for DRA allocated resources when configuring Pod and container cgroups, OOM scores, and Memory QoS thresholds.
    • Standard resource requests and limits can be resized in place for Pods that use DRA claims.
    • The scheduler supports unreferenced Pod-level claims, enforces resource limits with DRA, and enforces claim sharing rules, allowing claim sharing across Pods for overhead allocations.
    • Validation enforces constraints for the updated API and uses node declared features to confirm that target nodes have the NodeAllocatableDRA feature gate enabled. (#140009, @pravk03) [SIG API Machinery, Apps, Auth, Autoscaling, Node, Scheduling and Testing]
  • Updated the PodGroup API to replace PodGroupTemplateRef with WorkloadRef, a simpler and more direct way to reference the workload a PodGroup belongs to. (#140080, @dom4ha) [SIG API Machinery, Apps, Etcd, Node, Scheduling and Testing]
  • client-go: Removed nearly all context.TODO calls by introducing APIs where the caller passes in the context. Log calls use the logger provided by the caller when available. (#129125, @pohly) [SIG API Machinery, Apps, Architecture, Auth, CLI, Cloud Provider, Instrumentation, Network, Node, Storage and Testing]
  • kubelet: Added TLS support for gRPC container probes (behind a feature gate). (#137762, @amritansh1502) [SIG API Machinery, Apps, Node and Testing]

Feature

  • Added Beta support for compressed responses to WatchList requests. When a client sends Accept-Encoding: gzip, the API server returns a gzip compressed response. This behavior is enabled by default and can be disabled using the WatchListCompression feature gate. Regular watch requests are unaffected. (#140140, @p0lyn0mial) [SIG API Machinery]
  • Added GROUP, SCOPE, VERSIONS, and CREATED AT columns to kubectl get crd output, alongside the existing NAME column. The new columns provide additional information about the Custom Resource Definitions (CRDs) in a more concise and organized manner, making it easier for users to understand the details of their CRDs at a glance. (#131599, @jaehanbyun) [SIG API Machinery]
  • Added PodGroup methods to PodGroupManager and SharedLister, allowing scheduler plugins to obtain a consistent PodGroup state. (#140077, @macsko) [SIG Node, Scheduling and Testing]
  • Added Prometheus metrics for Windows kube-proxy (winkernel) load balancer operation failures: kubeproxy_sync_proxy_rules_winkernel_lb_create_failures_total, kubeproxy_sync_proxy_rules_winkernel_lb_update_failures_total, and kubeproxy_sync_proxy_rules_winkernel_lb_delete_failures_total. Each metric includes ip_family, lb_type, and error labels for fine-grained failure observability. (#137767, @princepereira) [SIG Instrumentation, Network and Windows]
  • Added ServiceName, PodManagementPolicy, and PersistentVolumeClaimRetentionPolicy to kubectl describe statefulset output. (#137547, @kfess) [SIG CLI]
  • Added AnnotatedEventf method to the new events API (EventRecorder and EventRecorderLogger interfaces in client-go/tools/events), enabling callers to attach custom annotations to events at creation time. (#138103, @adri1197) [SIG API Machinery and Node]
  • Added cpu_ids and memory fields at the pod level to the PodResources v1 API to report total allocated pod resources, while only returning container-level allocations for container-isolated containers. (#138738, @KevinTMtz) [SIG Node and Testing]
  • Added metrics.k8s.io/v1 support to kubectl top. (#139726, @tico88612) [SIG CLI and Instrumentation]
  • Added net.ipv4.tcp_slow_start_after_idle and net.ipv4.tcp_notsent_lowat to the allowed safe sysctls list. (#138389, @gheffern) [SIG Auth, Network and Node]
  • Added a --max-depth flag to kubectl explain --recursive to limit the depth of nested fields displayed in the output. (#138809, @shady0503) [SIG CLI and Testing]
  • Added a warning when kube-proxy is started without an explicitly specified proxy mode (such as iptables, ipvs, or nftables), because the default mode on Linux will switch from iptables to nftables in a future release. (#139957, @danwinship) [SIG Network]
  • Added an Alpha feature gate, ConsistentListFromCacheSkipTimeoutFallback. When enabled, kube-apiserver returns HTTP 429 for consistent LIST requests that cannot be served from the watch cache within the timeout window, instead of falling back to storage. (#138701, @yedou37) [SIG API Machinery]
  • Added an error outcome to the route_sync_total metric for failed route reconciles, alongside the existing changed and noop outcomes. (#140824, @lukasmetzner) [SIG Cloud Provider and Instrumentation]
  • Added context-aware APIs to client-go for REST mapping and discovery, replacing internal blocking API calls that used context.TODO. Consumers of client-go are encouraged to switch to the new APIs. The old APIs are marked as deprecated, but there is no plan to remove them. (#129109, @pohly) [SIG API Machinery, Apps, Auth, Instrumentation, Node and Testing]
  • Added metric apiserver_watch_cache_initialization_duration_seconds recording the duration of the most recent watch cache initialization, labeled by group and resource. (#138767, @Jefftree) [SIG API Machinery and Instrumentation]
  • Added metrics for informer activity in kube-apiserver. (#139968, @michaelasp) [SIG API Machinery and Testing]
  • Added progress reporting to StorageVersionMigration conditions, allowing users to see how many objects a migration has processed. (#138875, @michaelasp) [SIG API Machinery and Apps]
  • Added scheduler metrics for the topology-aware scheduling (TAS) placement phases, available when the TopologyAwareWorkloadScheduling feature gate is enabled: scheduler_generated_placements_total, scheduler_placement_evaluations_total, and scheduler_placement_evaluation_duration_seconds. (#139604, @alimaazamat) [SIG Instrumentation, Scheduling and Testing]
  • Added scheduler performance benchmark suites comparing preemption behavior across standalone (non-PodGroup) Pods and PodGroups with single and all disruption modes, and reorganized the default preemption performance tests into a dedicated directory. (#140651, @vshkrabkov) [SIG Scheduling and Testing]
  • Added structured CauseType values to PodDisruptionBudget-related eviction Forbidden errors in the eviction API, allowing clients to programmatically distinguish PDB invalid-state errors from other forbidden errors without string-matching on the message. (#138003, @shady0503) [SIG Apps, Auth and Node]
  • Added support for CBOR encoding in discovery endpoints and structured error responses when the CBORServingAndStorage feature gate is enabled. (#139632, @benluddy) [SIG API Machinery and Testing]
  • Added support for testing invariant metrics in integration tests. (#137883, @lalitc375) [SIG API Machinery, Auth and Testing]
  • Added support in client-go certificate manager Config for specifying a custom GenerateKey function. (#138999, @neolit123) [SIG API Machinery and Auth]
  • Added the 90s, 120s, 180s, and 300s buckets to the watch_list_duration_seconds metric. (#140757, @richabanker) [SIG API Machinery and Instrumentation]
  • Added the Alpha apiserver_watch_events_dispatch_duration_seconds metric, recording the duration from when a watch event is decoded from etcd until it is written to the watcher's outgoing result channel. (#140336, @richabanker) [SIG API Machinery, Etcd and Instrumentation]
  • Added the Alpha kubelet metric kubelet_pod_deferred_resize_duration_seconds histogram and the priority_bucket label on the kubelet_pod_pending_resizes gauge. (#140122, @natasha41575) [SIG Instrumentation and Node]
  • Added the Alpha kubelet metric pod_level_resources_admission_total to track adoption of Pod-Level Resources (KEP-2837) upon Pod admission, categorized by resource configuration mode and QoS class. (#140463, @ndixita) [SIG Instrumentation and Node]
  • Added the +k8s:dependentRequired("siblingJSONName") declarative validation tag. When the tagged field is set, the named sibling must also be set. (#139164, @yongruilin) [SIG API Machinery]
  • Added the --proxy-url flag to kubectl to override the proxy URL configured in the kubeconfig. (#139862, @Mujib-Ahasan) [SIG CLI]
  • Added the CompositePodGroup feature gate to enable Composite Pod Group functionality. (#139407, @jdzikowski) [SIG Scheduling]
  • Added the EtcdRangeStream beta feature gate. The watch cache initializes by streaming objects from etcd in a single RangeStream RPC instead of paginated Range requests. (#136915, @Jefftree) [SIG API Machinery, Architecture, Auth, CLI, Cloud Provider, Cluster Lifecycle, Etcd, Instrumentation, Network, Node, Scheduling, Storage and Testing]
  • Added the KubeProxyIPVS feature gate in preparation for deactivating and then removing the ipvs mode of kube-proxy. (#139397, @adrianmoisey) [SIG Network]
  • Added the PodGroup field to the PodGroupInfo object in kube-scheduler to enable plugins to obtain a consistent state throughout the scheduling cycle. (#140075, @macsko) [SIG Scheduling]
  • Added the WatchListCompression feature gate (Beta, enabled by default) to compress WatchList responses with gzip for clients that send Accept-Encoding: gzip. Regular Watch requests are unaffected. (#139308, @p0lyn0mial) [SIG API Machinery]
  • Added the allocatedPods kubelet endpoint, which surfaces the kubelet's allocated Pod spec for debugging in-place Pod resizing and other Pod update issues. Requires the KubeletAllocatedPodsEndpoint feature gate. (#140856, @tallclair) [SIG Node and Testing]
  • Added the cache_to_watcher stage to the Alpha apiserver_watch_events_dispatch_duration_seconds metric to measure the latency incurred when pushing events to a watcher's result channel. (#140851, @richabanker) [SIG API Machinery and Instrumentation]
  • Added the client-go informer metrics informer_store_resource_version, informer_queued_items, and informer_processing_latency_seconds to kube-scheduler, labelled name="kube-scheduler". (#140511, @Jefftree) [SIG Scheduling and Testing]
  • Added the incompletePodGroupPods data structure to the scheduling queue to store Pods waiting for their PodGroup object to be observed by kube-scheduler. (#139952, @macsko) [SIG Node, Scheduling and Testing]
  • Added the owner_api_group and owner_api_kind labels to the dynamic_resource_allocation_resourceclaim_creates_total metric to distinguish ResourceClaims created for Pods from those created for PodGroups under the DRAWorkloadResourceClaims feature gate. (#140422, @nojnhuh) [SIG API Machinery, Apps, Instrumentation, Node, Scheduling and Testing]
  • Added the queued_entities and queue_incoming_entities_total scheduler metrics. queued_entities reports the current number of scheduling entities (Pods or PodGroups) in scheduler queues, and queue_incoming_entities_total reports the total number of scheduling entities added to scheduler queues. (#139840, @iomarsayed) [SIG Instrumentation, Network and Scheduling]
  • Added the storage_to_cache stage to the Alpha apiserver_watch_events_dispatch_duration_seconds metric to track the latency from backend decode to watch cache ingestion. (#140860, @richabanker) [SIG API Machinery and Instrumentation]
  • Added the trigger (periodic or node_change) and outcome (changed or noop) labels to the Alpha route controller metric route_controller_route_sync_total. This lets operators observe how often periodic reconciliation corrects route drift versus running as a no-op. (#140147, @lukasmetzner) [SIG Cloud Provider and Instrumentation]
  • Added the scheduler extension point PlacementFeasible to allow early termination of the PodGroup scheduling cycle. This extension point is used by the GangScheduling plugin to stop evaluating pods once minCount becomes unsatisfiable. (#138643, @brejman) [SIG Scheduling and Testing]
  • Added the standard device attribute resource.kubernetes.io/numaNode and sysfs-based helper functions for DRA drivers. (#139929, @johnahull) [SIG Node]
  • Added validation to PodGroup scheduling that, when the PodGroupPreemptionPolicy feature gate is enabled, ensures that the preemption policies of Pods being evaluated for scheduling match the priority of the PodGroup. (#140359, @ania-borowiec) [SIG Scheduling and Testing]
  • Added validation to PodGroup scheduling which ensures priorities of the evaluated pods match the priority of the PodGroup. (#139920, @brejman) [SIG Scheduling and Testing]
  • Added workload preemption metrics in Alpha behind the WorkloadAwarePreemption feature gate. (#139373, @brejman) [SIG Instrumentation and Scheduling]
  • Applied --field-selector to pod metrics when invoking kubectl top pod. (#139107, @Mujib-Ahasan) [SIG CLI]
  • Changed PodGroup preemption to run after a failed PodGroup scheduling attempt for PodGroups with scheduling constraints. (#140683, @Argh4k) [SIG API Machinery, Scheduling and Testing]
  • Changed kube-apiserver with --enable-aggregator-routing=true to evenly load-balance requests across admission webhook endpoints, preventing connection caching from routing all concurrent requests to a single backend endpoint. Cluster administrators can temporarily opt out of this behavior using the WebhookRoundTripLoadBalancing feature gate (Beta, default true). (#139237, @aojea) [SIG API Machinery and Testing]
  • Changed the PatchPodStatus API in the scheduler framework to accept a slice of Pod conditions ([]*v1.PodCondition) instead of a single condition (*v1.PodCondition). This allows scheduler plugins to update multiple Pod conditions in a single API call, preventing newer calls from overwriting older ones when multiple conditions need to be updated concurrently. (#135160, @KunWuLuan) [SIG Scheduling]
  • Demoted the SchedulerPreQueueingHints feature gate from Beta to Alpha, disabled by default, because of issues found shortly before release. (#140959, @sanposhiho) [SIG Scheduling]
  • Enabled scaling to and from zero by default for the HorizontalPodAutoscaler (HPA). (#139648, @johanneswuerbach) [SIG Apps, Autoscaling and Testing]
  • Enabled the EtcdRangeStream feature gate by default and promoted it to Beta. (#140085, @Jefftree) [SIG API Machinery]
  • Enhanced Pod-by-Pod preemption to support PodGroups as preemption victims. (#137981, @vshkrabkov) [SIG Scheduling and Testing]
  • Updated pod group preemption errors to be prefixed with pod group preemption: message. (#139218, @Argh4k) [SIG Scheduling]
  • Functions and structs that take in authorizer.Authorizer might choose to accept only a smaller interface, authorizer.UnconditionalAuthorizer, in case only the receiver only needs to perform unconditional authorization requests and wants to signal this in the code for clarity. Any authorizer implementation must still implement the full authorizer.Authorizer interface. (#138801, @luxas) [SIG API Machinery, Auth, Node, Scheduling and Testing]
  • Graduated WatchCacheInitializationPostStartHook to GA. (#139452, @serathius) [SIG API Machinery]
  • Graduated the ConcurrentWatchObjectDecode feature gate to Beta, enabled by default. (#139679, @Jefftree) [SIG API Machinery and Etcd]
  • Graduated the ManifestBasedAdmissionControlConfig feature gate to Beta and enabled it by default. (#140559, @BenTheElder) [SIG API Machinery]
  • Graduated the NativeHistograms feature gate to Beta. (#140124, @richabanker) [SIG Architecture and Instrumentation]
  • Graduated the RelaxedServiceNameValidation feature gate to GA. (#139282, @adrianmoisey) [SIG Apps and Network]
  • Graduated the scheduler_plugin_execution_duration_seconds and scheduler_scheduling_algorithm_duration_seconds metrics from Alpha to Beta. (#138176, @abhay1999) [SIG Instrumentation, Scheduling and Testing]
  • Improved node health checks by verifying lease staleness with a live GET before marking nodes unhealthy, avoiding false positives from stale cache. (#138698, @michaelasp) [SIG Apps, Auth and Node]
  • Improved scheduling performance for required Pod affinity and anti-affinity with topologyKey: kubernetes.io/hostname, behind the InterPodAffinityHostnameFastPath feature gate. (#138198, @tetianakh) [SIG Apps, Node, Scheduling and Testing]
  • Made it possible for authorizers to return conditional decisions in addition to unconditional (Allow/Deny/NoOpinion). (#137204, @luxas) [SIG API Machinery, Auth, Node, Scheduling and Testing]
  • Optimized CEL admission policy evaluation by adopting a lazy zero-allocation reflection-based utility for object traversal, significantly reducing CPU usage and garbage collection overhead during request processing. (#138771, @lalitc375) [SIG API Machinery, Architecture, Auth, CLI, Cloud Provider, Cluster Lifecycle, Network, Node, Scheduling and Storage]
  • Optimized kube-scheduler performance for Pods with PersistentVolumeClaim mounts by processing only delta counts between scheduling cycles. (#139238, @yue9944882) [SIG Scheduling and Testing]
  • Preserved data in the DRA-related Pod status fields resourceClaimStatuses, extendedResourceClaimStatus, and nodeAllocatableResourceClaimStatuses when handling Pod status updates that omit those fields. This prevents updates from older clients from unsetting these DRA fields, which could leave Pods permanently stuck in Terminating. (#139876, @ashishpatel26) [SIG API Machinery, Auth, Node, Scheduling and Testing]
  • Promoted serviceaccount_legacy_tokens_total, serviceaccount_stale_tokens_total and serviceaccount_valid_tokens_total to Beta. (#137072, @tico88612) [SIG Auth, Instrumentation and Testing]
  • Promoted support for kubectl get -o kyaml to Stable. (#140076, @soltysh) [SIG CLI]
  • Promoted the AllowUnsafeMalformedObjectDeletion feature gate to Beta, enabled by default. List errors for objects that cannot be read from storage include the first underlying cause in the error message. (#140785, @ibihim) [SIG API Machinery and Etcd]
  • Promoted the DRAResourceClaimDeviceStatus feature gate to GA. (#137546, @LionelJouin) [SIG Node and Testing]
  • Promoted the InPlacePodVerticalScalingInitContainers feature gate to GA. (#140728, @natasha41575) [SIG Apps and Node]
  • Promoted the PLEGOnDemandRelist feature gate to GA. (#140805, @tallclair) [SIG Node]
  • Promoted the PodAndContainerStatsFromCRI feature gate to Beta, disabled by default. (#140081, @dgrisonnet) [SIG Node]
  • Promoted the PodLevelResourceManagers feature gate to Beta, enabled by default. (#140573, @KevinTMtz) [SIG Node and Testing]
  • Promoted the PodReadyToStartContainers condition to GA. (#140488, @Priyankasaggu11929) [SIG Node and Testing]
  • Promoted the kube-apiserver webhook metrics apiserver_webhooks_x509_missing_san_total and apiserver_webhooks_x509_insecure_sha1_total to Beta and updated their documentation. (#136894, @LoginovIlia) [SIG API Machinery, Instrumentation and Testing]
  • Promoted the kubelet PodsAPI gRPC service to Beta. (#140286, @briansonnenberg) [SIG Instrumentation, Node and Testing]
  • Reduced the scope of EventedPLEG to only accelerate detection of unexpected container terminations. (#139262, @HirazawaUi) [SIG Node]
  • Retried binding API calls in kube-scheduler when a transient error occurs. (#138855, @antekjb) [SIG Scheduling]
  • Set the KUBECTL_PATH environment variable to the path of the kubectl binary when it executes a plugin. (#138694, @brianpursley) [SIG CLI and Testing]
  • Set the nominatedNodeName field on pods from a PodGroup after a successful PodGroup preemption, consistent with single-pod preemption. (#138967, @antekjb) [SIG Scheduling and Testing]
  • The MaxUnavailableStatefulSet feature is enabled by default. (#139466, @soltysh) [SIG Apps]
  • Updated the apiserver_storage_list_* metrics to include storage and index labels to distinguish the storage backend and lookup path used to serve LIST requests. (#139125, @yedou37) [SIG API Machinery, Etcd and Instrumentation]
  • Updated the scheduler to avoid redundant preemption attempts during PodGroup scheduling when terminating victim pods are already present on the nominated nodes. (#138710, @mm4tt) [SIG Scheduling]
  • Added three different subtypes of the cluster event resource "Pod": "AssignedPod", "UnscheduledPod", "TargetPod". Plugins can and are expected to register to specific pod events for better performance. (#135905, @iomarsayed) [SIG Node, Scheduling, Storage and Testing]
  • Updated CoreDNS to v1.14.3. (#138536, @yashsingh74) [SIG Cloud Provider and Cluster Lifecycle]
  • Updated CoreDNS to v1.14.4. (#139735, @yashsingh74) [SIG Cloud Provider and Cluster Lifecycle]
  • Updated CoreDNS to v1.14.6. (#140497, @yashsingh74) [SIG Cloud Provider and Cluster Lifecycle]
  • Updated PodGroup scheduling to requeue remaining unscheduled Pods directly to the active queue (rather than backoff queue) after successful PodGroup scheduling, preserving their original timestamps so they retain scheduling precedence unless a higher priority entity is added. (#139613, @iomarsayed) [SIG Scheduling and Testing]
  • Updated PodGroup scheduling to skip PostFilter plugins for Pods in a PodGroup cycle. Instead, PodGroupPostFilter runs only when the entire PodGroup is unschedulable. (#140412, @Argh4k) [SIG Scheduling and Testing]
  • Updated PodGroup status to include the pod group preemption found a placement for podgroup, preempting <victim_count> victims message when workload-aware preemption finds a placement. (#140311, @Argh4k) [SIG Scheduling and Testing]
  • Updated admission webhooks to skip auth/authz virtual resources, such as tokenreviews and subjectaccessreviews, that ValidatingAdmissionPolicy and MutatingAdmissionPolicy already exclude. Added the ExcludeAdmissionWebhookVirtualResources feature gate in Beta, enabled by default, with an opt-out option to restore the previous behavior. (#140019, @BenTheElder) [SIG API Machinery and Testing]
  • Updated cri-tools to v1.36.0. (#138613, @saschagrunert) [SIG Cloud Provider and Node]
  • Updated default preemption to include the message preemption: found a potential placement for pod on node <node_name>, preempting <victim_count> victims in the FailedScheduling event and PodScheduled condition when it finds a potential Node for a Pod. (#140180, @Argh4k) [SIG Scheduling]
  • Updated the Go version used to build Kubernetes to v1.26.3. (#138864, @BenTheElder) [SIG Release]
  • Updated the Go version used to build Kubernetes to v1.26.4. (#139479, @BenTheElder) [SIG Release]
  • Updated the Go version used to build Kubernetes to v1.26.4. (#139584, @cpanato) [SIG Release]
  • Updated the Go version used to build Kubernetes to v1.26.5. (#140576, @palnabarun) [SIG Release]
  • Updated the Go version used to build Kubernetes to v1.26.6. (#141363, @BenTheElder) [SIG Release]
  • Updated the WorkloadAwarePreemption feature to perform a single scheduling attempt with all potential victims removed. This significantly improves performance but can result in a less optimal choice of preemption victims. (#139980, @Argh4k) [SIG Scheduling and Testing]
  • Updated volume mount host path type mismatch errors to log the actual path type alongside the expected one. (#121873, @skitt) [SIG Storage]
  • Updated workload-aware preemption to preempt victims so that as many as possible of the preemptor pods can be scheduled. (#138757, @jdzikowski) [SIG Scheduling and Testing]
  • kube-controller-manager: Deferred syncing an HPA object in the HPA controller when the controller has not yet observed HPA status writes from the last time the object was synced. (#139025, @omerap12) [SIG Apps and Autoscaling]
  • kube-scheduler: Added PlacementCycleState to the scheduling framework, providing per-placement state to PlacementScore plugins under the Alpha TopologyAwareWorkloadScheduling feature gate. (#138274, @wtravO) [SIG Scheduling]
  • kube-scheduler: Added support for PodGroups in the scheduling queue. The active, backoff, and unschedulable queues were abstracted to store QueuedEntityInfo (handling either individual pods or pod groups). (#138567, @macsko) [SIG Instrumentation, Scheduling and Testing]
  • kubeadm: Added the kubeproxydaemonset patch target to allow patching the kube-proxy DaemonSet during kubeadm init and kubeadm upgrade, consistent with the existing corednsdeployment patch target. (#138090, @SataQiu) [SIG Cluster Lifecycle]
  • kubeadm: Changed the preflight Port-xx checks for kube-apiserver, kube-scheduler, kube-controller-manager, and etcd to bind to the address configured in the kubeadm config for the respective component (via the localAPIEndpoint.address field or the --bind-address extraArgs override), instead of calling net.Listen() without an address (which binds to all available unicast and anycast IP addresses for the port). (#138250, @lentzi90) [SIG Cluster Lifecycle]
  • kubeadm: Removed the NodeLocalCRISocket feature gate which graduated to GA and was locked to enabled by default in a previous release. (#138645, @neolit123) [SIG Cluster Lifecycle]
  • kubeadm: The preflight check ContainerRuntimeVersion validates if the installed container runtime supports the RuntimeConfig gRPC method. For older kubelet versions than v1.38, it will return a preflight warning. (#139122, @carlory) [SIG Cluster Lifecycle]
  • kubelet: Deferred the deprecation removal timeline for the configuration flags (and the related fallback behavior) from v1.37 to v1.38 to align with containerd v1.7 support. (#139121, @carlory) [SIG Node and Testing]

Documentation

  • Corrected the kube-proxy --metrics-bind-address flag documentation by removing the incorrect statement that setting the flag to an empty string disables the metrics server. (#138940, @kairosci) [SIG Network]
  • Fixed a nil pointer dereference in client-go event key generation by adding nil checks in getEventKey, getSpamKey, and EventAggregatorByReasonFunc, preventing panics when processing nil events. (#135925, @jianzhangbjz) [SIG API Machinery]
  • Updated Japanese translation for kubectl. (#131176, @yude) [SIG CLI and Testing]
  • Updated client-go and apimachinery to track Go API changes in a Go-API/CHANGELOG.md file. (#138351, @pohly) [SIG API Machinery]
  • Updated the kubectl explain long description to fix formatting in the generated documentation and improve consistency with other commands. (#140357, @ingyeoking13) [SIG CLI]

Failing Test

  • Fixed a bug in kube-scheduler when the DRADeviceTaintRules feature gate is enabled that could cause scheduler panics when DeviceTaintRules exist and ResourceSlices change, or cause new DeviceTaintRules changes to be ignored. (#139651, @nojnhuh) [SIG Node and Testing]
  • Fixed a bug where nomination of a gated pod wasn't preventing lower-priority pods from scheduling on the nominated space. (#139057, @macsko) [SIG Scheduling]

Bug or Regression

  • Added apiserver_storage_list_duration_seconds, a metric measuring end-to-end apiserver list latency (etcd read plus object decode), labelled by whether etcd RangeStream was used, so streamed and non-streamed lists can be compared directly. (#140697, @Jefftree) [SIG API Machinery, Etcd and Instrumentation]
  • Added the HPAOptimizedSelectorStore feature gate in Beta, enabled by default, to reduce lock contention in the HorizontalPodAutoscaler controller's selector overlap detection and improve reconciliation throughput at high HorizontalPodAutoscaler counts and concurrency. (#139142, @hakuna-matatah) [SIG Apps and Autoscaling]
  • Added the group name to the kubectl error message when a resource type is not found under the specified group, for example the server doesn't have a resource type "pdb" in group "hpa". (#140759, @makeittotop) [SIG CLI]
  • Avoided costly comparisons during SELinux metric emission. (#138981, @gnufied) [SIG Apps and Storage]
  • Changed client-go RetryWatcher to log 410 Gone (resource expired) errors at debug verbosity (V(4)) instead of ERROR level during watch establishment. (#138295, @kencochrane) [SIG API Machinery]
  • Changed kube-apiserver to validate the --advertise-address IP when using --endpoint-reconciler-type master-count or lease, ensuring the specified IP address can be persisted to an Endpoints API object successfully. (#138102, @kairosci) [SIG API Machinery]
  • Changed kube-proxy to skip full-sync operations when operating in large-cluster mode (more than 1000 endpoints). (#138571, @aojea) [SIG Network]
  • Changed kube-proxy to truncate nftables comments to the kernel's 128-byte limit before programming service maps, avoiding sync failures for long Service names. (#139516, @Vinayak9769) [SIG Network]
  • Changed kubectl get to return an error when --label-columns is used with custom-columns output. (#138094, @ahmadmaha02) [SIG CLI]
  • Changed image volume validation to reject empty image.reference fields in Pod templates (Deployment, StatefulSet, DaemonSet, Job, etc.). (#135989, @Okabe-Junya) [SIG Apps and Node]
  • Changed the HPA controller to reconcile newly created and spec-changed HPAs immediately instead of waiting for the full resync period (default 15s). (#138294, @Fedosin) [SIG Apps and Autoscaling]
  • Changed the kubelet to enforce explicit HTTP method restrictions for logs-related endpoints. Read-only kubelet server endpoints reject non-GET methods with 405. NodeLogQuery allows only GET and POST and rejects other methods with 405. (#138088, @amritansh1502) [SIG Node]
  • DRA: Fixed a bug where a rare missed informer update of a ResourceClaim could cause Pods to remain pending until the unschedulable queue was flushed. (#140831, @pohly) [SIG Node and Scheduling]
  • Disabled the PodLevelResourceManagers feature gate by default because of critical issues found before release. (#141209, @SergeyKanzhelev) [SIG Node]
  • Enabled evaluation of list-type attributes, the .includes function, and CEL macros in kube-scheduler even when the ListTypeAttributes feature gate is disabled. This prevents errors during rolling upgrades or when the feature gate is toggled. (#139395, @everpeace) [SIG Node]
  • Enabled netlink support by default in kube-proxy nftables mode. kube-proxy uses netlink directly to list rules and chains, improving performance by avoiding running and parsing the nft command-line binary. The NFTablesNetlink feature gate (Beta, enabled by default) can be disabled to restore the previous behavior. (#137536, @aojea) [SIG Network and Testing]
  • Fixed 409 Conflict errors between the PVC protection controller and the PV binder during initial PVC binding. The Unused condition is evaluated only after the PersistentVolumeClaim is bound. (#140833, @huww98) [SIG Apps and Storage]
  • Fixed CEL behavior for set and map lists. Equality (==) no longer matches lists containing duplicates, and concatenation (+) correctly applies set and map merge semantics to appended elements. (#140293, @jpbetz) [SIG API Machinery]
  • Fixed DRA scheduling bugs where the structured allocator incorrectly counted a device's shared counters while evaluating candidates. It could retain a counter reservation after rejecting or backtracking from a candidate, or clear a shared device's in-use marker, causing a later allocation to charge the counter twice. These issues could cause the allocator to incorrectly treat a counter set as exhausted and leave a Pod pending on a node that could satisfy its resource requirements. This affected the allocator used by the default feature configuration. (#140431, @thc1006) [SIG Node]
  • Fixed Pod-level MemoryQoS memory protection (memory.min and memory.low) being silently dropped during in-place Pod resize, and added Pod-level memory.high enforcement when the PodLevelResources feature gate is enabled. (#140262, @sohankunkerkar) [SIG Node and Testing]
  • Fixed VolumeAttachment validation to report the correct maximum message size (1024 bytes) in error messages. (#136436, @Okabe-Junya) [SIG Storage]
  • Fixed Windows CPU affinity so that, when the CPU Manager static policy and the Memory Manager are both active under the WindowsCPUAndMemoryAffinity feature gate, containers are pinned to the CPU Manager's allocated set instead of the union of that set and every CPU on the selected NUMA node. The previous behavior could overlap CPUs exclusively allocated to other containers. (#139684, @zylxjtu) [SIG Node]
  • Fixed DecodeMetadataFromStream to skip only entries with unknown API versions and return errors for decode failures or malformed metadata in supported versions, preventing silent data loss. (#138530, @alaypatel07) [SIG Node]
  • Fixed kube-apiserver hanging indefinitely on SIGTERM when it could not create its identity Lease, for example, because the hostname exceeded 63 bytes. (#140241, @camilamacedo86) [SIG API Machinery]
  • Fixed kube-proxy to remove stale conntrack entries when a UDP Service no longer has any serving endpoints, for example after scaling to zero. This prevents previously established one-way UDP flows from being indefinitely blackholed to deleted Pod IPs. (#139629, @Bafff) [SIG Network]
  • Fixed kubectl cluster-info dump --output-directory creating world-readable dump files. Files are created with mode 0600 and kubectl-created directories with mode 0700, because dumped Pod logs can contain sensitive data. (#140189, @ashvinctrl) [SIG CLI and Security]
  • Fixed kubectl get storageclass to show only the effective default StorageClass as "(default)" when multiple StorageClasses have the default annotation. (#135964, @jaehanbyun) [SIG CLI and Storage]
  • Fixed kubelet applying device health updates to the wrong Pod status when device plugins for different resources exposed devices with identical IDs. Affected Pods reflect device health changes immediately instead of waiting for the next periodic Pod sync. (#140323, @harche) [SIG Node]
  • Fixed kubelet failure starting on ZFS due to missing cadvisor plugin. (#138587, @BenTheElder) [SIG Node]
  • Fixed a DRA consumable-capacity scheduling bug where a device that consumes shared counters could have them counted twice when it already had a persisted shared allocation with no consumed capacity, wrongly rejecting a later claim for the same device and leaving the Pod pending. (#140437, @thc1006) [SIG Node and Scheduling]
  • Fixed a DRA issue where drivers might not recreate ResourceSlices that were deleted externally, depending on timing between driver updates and deletions. (#140063, @pohly) [SIG API Machinery, Apps, Node and Testing]
  • Fixed a DRA partitionable devices issue where counters published by a DRA driver outside the valid int64 range could be mutated in the informer cache, causing future Pod allocations to fail incorrectly. (#140518, @weizhoublue) [SIG Node]
  • Fixed a DRA scheduling bug where the structured allocator keyed shared-counter caches by pool name only, causing two drivers with the same pool name on a node to use each other's counter definitions and incorrectly accept or reject device allocations in the second driver's pool. (#140435, @thc1006) [SIG Node]
  • Fixed a Dynamic Resource Allocation (DRA) scheduler bug that could assign mutually exclusive device partitions to multiple Pods. This affected DRA drivers using SharedCounters (DRAPartitionableDevices) together with multi-allocatable devices (DRAConsumableCapacity). Depending on the device and driver, the incorrect double-allocation could cause workload failures, device conflicts, crashes, or data loss. (#139040, @ashvindeodhar) [SIG Node]
  • Fixed a kube-proxy IPVS-mode performance bug where syncProxyRules could take tens of seconds in clusters with many Services because GetAllLocalAddressesExcept issued one full netlink address dump per interface. The function issues a single dump per address family, reducing syncProxyRules latency by orders of magnitude on large clusters. (#138927, @ytcisme) [SIG Network]
  • Fixed a kube-proxy issue on Windows where transient HNS downtime during restart or recovery could cause incorrect LoadBalancer state reconciliation, resulting in duplicate LoadBalancer creation failures with Cannot create a file when that file already exists. (0xb7) errors. (#139503, @princepereira) [SIG Network and Windows]
  • Fixed a kubelet bug where init containers could be skipped when a Pod sandbox was recreated while a main container from the previous sandbox was still known to the container runtime. (#138514, @chez-shanpu) [SIG Node]
  • Fixed a kubelet issue where Pods with subPath mounts could become stuck in an error loop after FUSE or GlusterFS network filesystem disruptions. (#139275, @yuehaii) [SIG Node, Storage and Testing]
  • Fixed a kubelet memory leak regression in v1.36 caused by leaked contexts on every Pod sync. (#139850, @compumike) [SIG Node]
  • Fixed a kubelet panic in image pull credential verification when maxParallelImagePulls is configured above 31. (#138937, @RajvardhanPatil07) [SIG Node]
  • Fixed a v1.33 regression that could cause a panic in the endpoint controller when processing services with empty IPFamilies field (pre-dual-stack services that were never spec-updated). (#138736, @rahulbabu95) [SIG Apps and Network]
  • Fixed a v1.35 regression where exec readiness probes stopped executing (failing with "context canceled") once a pod began graceful termination, leaving the pod's Ready condition frozen during shutdown. (#140882, @karlkfi) [SIG Node]
  • Fixed a bug in CEL where quantity.Add mutated the receiver. (#140556, @jpbetz) [SIG API Machinery]
  • Fixed a bug in ImageLocality scoring where image volumes could receive a higher score than equivalent regular container images. (#138951, @sujoshua) [SIG Scheduling]
  • Fixed a bug in kube-apiserver where a request matching multiple ValidatingAdmissionPolicy bindings with audit actions only recorded the first validation failure in the audit annotation; all audit failures are published in a single annotation. (#140001, @lalitc375) [SIG API Machinery]
  • Fixed a bug in the DRA kubelet plugin helper where drivers with names longer than ~30 characters could not enable rolling updates because the plugin registration socket path exceeded the AF_UNIX path length limit. Rolling-update registration sockets pick the shortest basename that fits under the configured registry directory, preferring <driver>-<Pod UID>-reg.sock, then <driver>-<hashed UID>-reg.sock, then dra-<hashed driver+UID>-reg.sock. (#139623, @vishalanarase) [SIG Node and Testing]
  • Fixed a bug that caused Pods in a PodGroup sharing a ResourceClaim to get stuck scheduling. (#140269, @nojnhuh) [SIG Node, Scheduling and Testing]
  • Fixed a bug that could cause the admission controller to panic when evaluating CEL expressions against typed map lists with three keys. (#140386, @weizhoublue) [SIG API Machinery]
  • Fixed a bug when the GenericWorkload feature gate is enabled that could prevent Pods in the same PodGroup sharing the same ResourceClaim from successfully scheduling. (#139418, @nojnhuh) [SIG Node and Scheduling]
  • Fixed a bug where Burstable Pod memory.low soft protection was ineffective because the parent cgroup lacked the ancestor coverage required by the kernel's hierarchical protection model. (#140267, @sohankunkerkar) [SIG Node and Testing]
  • Fixed a bug where Pod .status.resourceClaimStatuses could flap between partial lists of claims when multiple claims were used in the Pod. (#138408, @johnbelamaric) [SIG Apps and Node]
  • Fixed a bug where Pods in a PodGroup sharing a ResourceClaim could be scheduled to Nodes where the ResourceClaim is not available. (#140089, @nojnhuh) [SIG Node, Scheduling and Testing]
  • Fixed a bug where Pods that share multi-node claims and also have per-node claims can get stuck in Pending. (#139017, @johnbelamaric) [SIG Node and Scheduling]
  • Fixed a bug where ResourceClaims using allocationMode: All with consumable capacity could be partially allocated when a matching device had insufficient remaining capacity. Such claims correctly fail to allocate until all matching devices can be satisfied. (#140769, @kiarashazarnia) [SIG Node]
  • Fixed a bug where ValidatingAdmissionPolicy and MutatingAdmissionPolicy evaluation could observe subtle differences (particularly around quantity fields and type meta fields) in object representation. The bug could occur when a parameter object was loaded via an alternative path, due to a cache miss. CEL expressions behave deterministically against params. (#140201, @jpbetz) [SIG API Machinery]
  • Fixed a bug where kubectl drain --disable-eviction --dry-run=server hangs indefinitely. (#137543, @kfess) [SIG CLI and Testing]
  • Fixed a bug where a StatefulSet with the OnDelete update strategy never updated Status.CurrentRevision to match Status.UpdateRevision after all Pods were recreated with the new revision. (#136833, @zhijun42) [SIG Apps]
  • Fixed a bug where disabling the MemoryQoS feature gate did not clear per-container memory.high cgroup values, causing containers to remain throttled at stale limits. (#139377, @sohankunkerkar) [SIG Node and Testing]
  • Fixed a bug where enabling the DRAListTypeAttributes feature gate could prevent device allocation even when a valid combination existed. This occurred when the allocator needed to backtrack while allocating multiple devices with a matchAttribute constraint using list-type attribute values. (#140325, @everpeace) [SIG Node and Scheduling]
  • Fixed a bug where kubelet would generate an event once per second for every image volume in a pod. (#138655, @mdbooth) [SIG Node]
  • Fixed a bug where non-admitted Pods could briefly count against the allocated budget, causing spurious admission failures for reasonably sized Pods. (#139522, @haircommander) [SIG Node]
  • Fixed a bug where pods with multiple subPath volume mounts on Windows would get stuck in Terminating state because file handles from subPath preparation were leaked, preventing volume cleanup. (#138367, @timmy-wright) [SIG Node, Testing and Windows]
  • Fixed a bug where successfully scheduled Pods could be stuck with the PodScheduled=False condition. (#139602, @nojnhuh) [SIG Node and Scheduling]
  • Fixed a bug where the kubelet node shutdown manager could leak D-Bus connections on repeated failures, eventually leading to thread exhaustion and crashes. (#137141, @harche) [SIG Node]
  • Fixed a bug where the kubelet did not enforce per-container ephemeral-storage limits on restartable init containers (sidecar containers), allowing them to exceed their declared limit without triggering pod eviction. (#138462, @shachartal) [SIG Node and Testing]
  • Fixed a case where Pods in a PodGroup that were successfully evaluated during a failed PodGroup scheduling cycle had nominatedNodeName set from that evaluation instead of from PodGroup preemption. (#140590, @Argh4k) [SIG Scheduling]
  • Fixed a concurrent map read/write data race in handleSchedulingFailure during scheduling failure handling. (#140623, @SparshGarg999) [SIG Scheduling]
  • Fixed a kube-scheduler panic when a DRA ResourceClaim using allocationMode: All selects a device that consumes shared counters. (#138885, @takonomura) [SIG Node]
  • Fixed a metrics leak in the scheduler PriorityQueue by decrementing UnschedulableReason metrics when Pods were deleted from active or backoff queues, ensuring accurate reporting of unschedulable plugin reasons. (#138482, @vshkrabkov) [SIG Scheduling]
  • Fixed a panic caused by integer division by zero and incorrect ResourceSlice admission validation when a DRA consumable capacity validRange step, min, max, or default value is negative or greater than 9223372036854775807. (#140666, @thc1006) [SIG Node]
  • Fixed a panic in ResourceSlice validation that could occur when the DRAConsumableCapacity feature gate was enabled and a capacity request policy set validRange.step to 0. (#139698, @wilmerdooley) [SIG Node]
  • Fixed a panic in kube-controller-manager that could occur when a StorageVersionMigration targeted a resource missing from the RESTMapper, for example, when a CRD was deleted while its migration was pending. (#140586, @zwindler) [SIG API Machinery and Apps]
  • Fixed a race condition in preemption, where a preemptor pod could get stuck in unschedulable state. (#139162, @brejman) [SIG Scheduling and Testing]
  • Fixed a race in kubelet where PrepareResources could attach a Pod to a ResourceClaim that was concurrently being unprepared, leaving the Pod running with unprepared devices. (#140527, @bart0sh) [SIG Node]
  • Fixed a regression in Kubernetes v1.35 where, with a Parallel Pod management policy, unavailable Pods from an older revision were incorrectly counted toward the maxUnavailable budget. (#137666, @soltysh) [SIG Apps]
  • Fixed a regression in Server-Side Apply where patching a container type (list or map) could return 422 required errors for apply requests that previously succeeded. (#140294, @jpbetz) [SIG API Machinery, Architecture, Auth, CLI, Cloud Provider, Cluster Lifecycle, Instrumentation, Network, Node, Scheduling, Storage and Testing]
  • Fixed a regression in v1.36 where modifications to scheduling directives (nodeSelector, tolerations, nodeAffinity) on suspended Jobs were rejected if the JobSuspended condition had not yet been set by the job controller. (#139287, @kannon92) [SIG Apps and Testing]
  • Fixed a regression in retrying deferred resizes caused by changes to Pod resource footprint calculation. (#140646, @natasha41575) [SIG Node and Scheduling]
  • Fixed a regression where the Job controller could report status.active as 0 while replacement Pod creation was deferred due to pod-failure backoff, causing the Job status update to be rejected by kube-apiserver. This could delay flushing uncounted terminated Pods, finalizer removal, and Job status updates, leaving Pods stuck in Terminating and the Job with stale status until the backoff elapsed. (#139457, @akhilsingh-git) [SIG Apps]
  • Fixed a regression where the kubelet did not clear stale cgroup v2 memory.min and memory.low values when the MemoryQoS feature gate was disabled after being previously enabled. (#138903, @sohankunkerkar) [SIG Node and Testing]
  • Fixed a scheduler bug in DRA consumable capacity where a ResourceSlice with a device capacity requirement stored as a high-precision decimal, such as a fine-grained fractional value or a value above the int64 range, could be mutated in the informer cache during allocation. This mutation could incorrectly cause allocation to fail for subsequent Pods. (#140702, @weizhoublue) [SIG Node]
  • Fixed a scheduler bug where clearing NominatedNodeName could leave Pods tracked under an empty node key in the scheduler's nominator. (#139904, @pacoxu) [SIG Scheduling]
  • Fixed a scheduler cache bug where assumed Pods were not removed correctly from PodGroupState after receiving a deletion timestamp. (#138445, @iomarsayed) [SIG Scheduling]
  • Fixed admission handling so that updates to namespaced objects that still exist after their namespace was deleted are allowed. (#140661, @sanchezl) [SIG API Machinery and Testing]
  • Fixed an issue in the CronJob controller where it failed to adopt existing Jobs by erroneously using the empty namespace from the jobTemplate. (#136920, @ysam12345) [SIG Apps]
  • Fixed an issue that could cause duplicate configuration entries to be reported in ResourceClaim status. (#139732, @LionelJouin) [SIG Node]
  • Fixed an issue where PodGroup preemption that detected an ongoing preemption would clear nominatedNodeName on the PodGroup's Pods. (#140641, @Argh4k) [SIG Scheduling and Testing]
  • Fixed an issue where the StatefulSet controller's skip metrics were not properly registered. (#138451, @michaelasp) [SIG API Machinery, Apps and Testing]
  • Fixed an issue where the kubelet would delete the CSI mount directory when a periodic NodePublishVolume call (triggered by setting CSIDriver.spec.requiresRepublish to true) returned an error, leaving the Pod with stale volume contents that subsequent successful republishes could not repair. (#139045, @aramase) [SIG Storage]
  • Fixed audit logging of malformed patch request bodies. (#139419, @hoskeri) [SIG API Machinery, Auth and Testing]
  • Fixed capacity accounting in the DRA consumable-capacity allocator. A capacity request that a device's integer or milli-value range arithmetic cannot represent, or that resolves to a negative value, is rejected instead of being treated as satisfiable and wrongly allocating a device or capping it to a smaller value. (#140442, @thc1006) [SIG Node]
  • Fixed duplicate logs when trying to attach to a pod fails. (#139091, @olamilekan000) [SIG CLI]
  • Fixed duplicated mount arguments in log string output from MakeMountArgsSensitiveWithMountFlags. (#138098, @jeffbearer) [SIG Storage]
  • Fixed handling of a certificate authority path outside the .kube/config directory on Windows, which was converted to a relative path instead of matching the behavior on other operating systems. (#135735, @bliles) [SIG API Machinery]
  • Fixed inconsistent ephemeral-storage formatting between capacity and allocatable values in Node status by using the DecimalSI format for ephemeral-storage capacity. (#137652, @0xMH) [SIG Node]
  • Fixed incorrect error message formatting in the HPA controller when object metric retrieval fails. Error messages correctly display the metric name, object kind, namespace, object name, and the underlying error. Also improved error wrapping across the HPA controller to use %w instead of %v, enabling proper error chain inspection. (#139029, @Fedosin) [SIG Apps and Autoscaling]
  • Fixed inter-pod affinity, anti-affinity, and volume restriction evaluation in kube-scheduler during PodGroup scheduling cycles. The scheduler snapshot's AssumePod and ForgetPod methods correctly maintain affinity node lists and PVC usage tracking. (#139054, @net0pyr) [SIG Scheduling and Testing]
  • Fixed nil pointer dereference in Windows memory eviction threshold notifier when GetPerformanceInfo() fails. (#138727, @rzlink) [SIG Node]
  • Fixed queue hint for inter-pod anti-affinity in case there are multiple terms, which might have caused delays in scheduling. (#139161, @brejman) [SIG Scheduling]
  • Fixed regression in kubectl resource printing on bigger data sets (100+ rows). (#138550, @rawkode) [SIG CLI]
  • Fixed stale remote HNS endpoint cleanup on Windows when a pod IP is reused across nodes in L2Bridge networks, preventing DNS timeouts caused by traffic being routed to the wrong node. (#138000, @princepereira) [SIG Network and Windows]
  • Fixed the DRA kubelet plugin helper repeating the listen error instead of reporting why removing a stale Unix domain socket failed when it could not start its listener. (#141105, @thc1006) [SIG Node]
  • Fixed the ResourceClaim controller mutating the shared informer cache when creating a ResourceClaim from a ResourceClaimTemplate that has annotations. (#141104, @thc1006) [SIG Apps and Node]
  • Fixed the kube-apiserver to create metadata fields for create-via-update and created-via-apply requests like they are for create requests. UID and resourceVersion preconditions are still honored. (#138908, @jpbetz) [SIG API Machinery and Testing]
  • Fixed the build for test/images/glibc-dns-testing. (#138877, @BenTheElder) [SIG Network and Testing]
  • Fixed the error message from PodGroupPostFilter to contain the correct extension point name. (#140747, @Argh4k) [SIG Scheduling]
  • Fixed the inconsistency between opportunistic batching and PodGroups that made the batching hints always infeasible during PodGroup scheduling cycle. (#138754, @macsko) [SIG Scheduling]
  • Fixed the wrong cause of the UnexpectedJob event/warning by checking the owner reference of the job correctly in the cron job controller. (#133313, @kei01234kei) [SIG Apps]
  • Generated metadata.generation and status.observedGeneration fields in HorizontalPodAutoscaler resources. (#138228, @adrianmoisey) [SIG API Machinery, Apps, Autoscaling and Testing]
  • Improved kubeadm join reliability by using the KubernetesAPICall timeout (default 1 minute) when fetching the kubeadm-config ConfigMap from the cluster, instead of the short 350ms retry previously used for optional component configs. A new shortConfigMapGet parameter was added to FetchInitConfigurationFromCluster so that callers like kubeadm reset can still use the short retry. (#139667, @damdo) [SIG Cluster Lifecycle]
  • Improved error reporting when invoking kubectl exec. (#138214, @hunshcn) [SIG CLI and Testing]
  • Improved scheduler handling of large PodGroups by reducing the likelihood of scheduling stalls when member Pods transiently fail to bind to Nodes, such as when many Pods share the same ResourceClaim. (#140478, @nojnhuh) [SIG Scheduling]
  • Improved the logic in kubeadm around warnings when a user sets a non-default bindAddress in KubeProxyConfiguration. (#139989, @vinayakray19) [SIG Cluster Lifecycle]
  • Improved the resilience of kubeadm etcd learner promotion by correctly handling cases where promotion succeeds but a transient client-side error is returned, preventing unnecessary etcd-join failures. (#139842, @jihyun-huh) [SIG Cluster Lifecycle]
  • Fixed kubelet to recover from corrupted subpath mount points (for example, stale NFS file handle) during container restart instead of leaving the pod stuck in CreateContainerConfigError. (#138856, @RomanBednar) [SIG Storage]
  • Removed an edge case that could allow malformed object deletion to bypass admission and graceful deletion of well-formed objects. (#137582, @benluddy) [SIG API Machinery, Etcd and Testing]
  • Removed the Alpha admission plugin that validated that PodGroup resources reference an existing Workload and match the declared PodGroupTemplate spec. (#139008, @wojtek-t) [SIG API Machinery, Etcd, Scheduling and Testing]
  • Reverted the cri-api KeyValue value field to its pre-v1.34 JSON encoding behavior for compatibility with earlier releases. (#139964, @liggitt) [SIG Node]
  • Surfaced the error reason when invalid service CIDRs are configured. (#139182, @PseudoResonance) [SIG Network]
  • Updated client-go to preserve the source file's permissions when migrating a kubeconfig file (for example, ~/.kube/.kubeconfig to ~/.kube/config). Previously, the destination file was created with broader permissions, which could expose credentials to other users on the same system. (#138142, @brianpursley) [SIG API Machinery]
  • Updated kube-proxy to exit when the watched Node's IPs change or when the Node object is deleted, allowing it to restart with updated node networking state. (#138183, @abishekgiri) [SIG Network]
  • Updated kubectl run error messages for invalid --restart and --image-pull-policy values to list the accepted values. (#138188, @ogormans-deptstack) [SIG CLI]
  • Updated kubelet eviction manager to exclude hugepage-reserved RAM from AvailableBytes in the calculation of memory.available on nodes with hugepages. This fixes delayed eviction and OOM kills caused by inflated available memory reporting. The HugepageAwareEviction feature gate (default: enabled) can be disabled to restore the previous behavior. (#138127, @jingczhang) [SIG Node and Testing]
  • Updated the PodGroup status.conditions field to reflect the failure reason when scheduling is rejected due to mismatched .spec.schedulerName across Pods in a group. (#140183, @Argh4k) [SIG Scheduling]
  • Updated the PodReadyToStartContainers condition to include a diagnostic message when status is False, explaining why the pod sandbox is not ready (for example, pod sandbox has no IP address or no pod sandbox exists). This improves debuggability for Pods stuck in ContainerCreating without requiring access to node logs. (#135300, @harche) [SIG Node]
  • Updated the kubelet to emit FailedToRetrieveImagePullSecret events only when an image pull has failed. (#138432, @Jamstah) [SIG Node]
  • Updated the kubelet to no longer emit V(4) "Label not found" logs for missing optional container annotations. (#140163, @HirazawaUi) [SIG Node]
  • Updated the pods/binding subresource endpoint to validate the specified node name consistently. (#136776, @yakir-shriker) [SIG Apps and Scheduling]
  • Updated the version of the nft binary in the kube-proxy image to nftables v1.0.6.1 to fix issues resyncing kube-proxy in nftables mode on systems containing rules created by recent versions of nftables. (#140405, @danwinship) [SIG Testing]
  • Used a stable curl download for the Windows busybox testing image. (#138879, @BenTheElder) [SIG Testing and Windows]
  • client-go: Added support for waiting for in-progress event handler runs to complete before closing the event handler. (#139755, @michaelasp) [SIG API Machinery]
  • client-go: Added the Bookmark and LastStoreSyncResourceVersion methods to FakeCustomStore so that it satisfies the cache.Store interface again, after those methods were added to the interface in v0.36. (#140966, @alancaldelas) [SIG API Machinery]
  • kubeadm: Changed kubeadm init so that, when the default admin.conf and super-admin.conf paths are used, the files are loaded but in-memory kubeconfigs are constructed pointing to InitConfiguration.localAPIEndpoint instead of ClusterConfiguration.controlPlaneEndpoint. This resolved issues with delayed load balancers that are provisioned only after the first kube-apiserver instance starts. (#138449, @neolit123) [SIG Cluster Lifecycle]
  • kubeadm: Changed kubeadm join to return a clear error message when the TLS bootstrap kubeconfig has a current-context that does not appear in the contexts list, instead of panicking with a nil pointer dereference. (#138853, @alexmchughdev) [SIG Cluster Lifecycle]
  • kubeadm: Changed cluster-info discovery over HTTPS to check the HTTP response status code, so a non-200 response produces a clear error instead of a confusing kubeconfig parse failure. (#138852, @alexmchughdev) [SIG Cluster Lifecycle]
  • kubeadm: Changed the etcd cluster status check to use a quorum approach instead of considering the health of all members, so the check no longer fails when there are sufficient healthy voting members. (#138403, @ahrtr) [SIG Cluster Lifecycle]
  • kubeadm: Fixed MemberPromote to skip the etcd promote API call when the member is already a voting member, avoiding unnecessary retries and timeout. (#138390, @wgkingk) [SIG Cluster Lifecycle]
  • kubeadm: Fixed a panic in kubeadm PKI key loading when the private key type and public key type mismatch. (#138939, @SataQiu) [SIG Cluster Lifecycle]
  • kubeadm: Fixed kubeadm init phase certs --dry-run to correctly copy existing CA files. (#139339, @ErikJiang) [SIG Cluster Lifecycle]
  • kubeadm: Skipped LocalAPIEndpoint defaulting on kubeadm join for worker nodes. (#138692, @clwluvw) [SIG Cluster Lifecycle]
  • kubeadm: Used a dedicated ClusterRole system:kubelet-api-admin for the kube-apiserver kubelet client. (#138957, @neolit123) [SIG Cluster Lifecycle]
  • kubelet/DRA: Fixed a bug where retrying a partially failed PrepareResources caused duplicate CDI device IDs to be passed to the CRI runtime, which could cause container start to fail. (#140274, @bart0sh) [SIG Node]
  • kubelet: Changed the DefaultPodSysctls feature to treat an unset spec.hostUsers as true when evaluating user.* sysctls. (#140892, @weizhoublue) [SIG Node]
  • kubelet: Fixed a goroutine leak on shutdown by making the eviction manager's monitoring goroutine exit promptly when the kubelet context is cancelled. (#138854, @alexmchughdev) [SIG Node]
  • kubelet: Fixed incorrect Pod-level CPU requests reported in status from the cgroup v2 readback. (#137660, @pacoxu) [SIG Node]
  • kubelet: Populated involvedObject.uid on node events on a best-effort basis once the node is registered, so node events can be correlated by UID, for example in kubectl describe node. The UID is resolved once and not refreshed afterward. If a node is deleted and recreated with a new UID while the kubelet keeps running, its events continue to use the original UID until the kubelet restarts. (#139921, @harche) [SIG Node]
  • kubelet: Set cgroup v2 memory.high for BestEffort containers when MemoryQoS is enabled (per KEP-2570). (#138139, @amritansh1502) [SIG Node]

Other (Cleanup or Flake)

  • Added the HasValidationFunc method to runtime.Scheme to report whether a declarative validation function is registered for a type. (#140120, @yongruilin) [SIG API Machinery and Testing]
  • Changed MutatingAdmissionPolicy and MutatingAdmissionPolicyBinding storage in etcd to use the admissionregistration.k8s.io/v1 API version. (#137375, @Jefftree) [SIG API Machinery, Etcd and Testing]
  • Changed ResourceClaim config status to leave the requests field empty when the configuration applies to all requests. (#139731, @LionelJouin) [SIG Node]
  • Changed client-go to request v2 for aggregated discovery instead of falling back to v2beta1. (#138271, @Jefftree) [SIG API Machinery]
  • Changed the kube-apiserver service/proxy subresource to use EndpointSlices instead of Endpoints when proxying to a Service. This change only affects clusters that manually create Endpoints for a Service and have EndpointSlice mirroring disabled. (#134860, @danwinship) [SIG API Machinery and Network]
  • Changed the scheduler's opportunistic batching to rescore the previously chosen node when it is still feasible, allowing it to compete with cached candidates for the next hint rather than always being skipped. (#140289, @romanbaron) [SIG API Machinery, Etcd, Instrumentation, Scheduling and Testing]
  • DRA: Fixed a potential crash in the scheduler, recovered after restart, when the ResourceSlice tracker encountered an OnDelete event for a DeviceTaintRule whose deleted object is unknown. (#140193, @pohly) [SIG API Machinery and Node]
  • DRA: Locked the DRAPrioritizedList feature gate to enabled by default. The Prioritized List feature reached GA in v1.36 and can no longer be disabled. (#139110, @mortent) [SIG Node, Scheduling and Testing]
  • Deprecated MultiLock, UnknownLeader, and ConcatRawRecord in client-go leader election resourcelock package. (#138070, @Jefftree) [SIG API Machinery]
  • Fixed a bug in kubelet DRA where deleting a Pod could unprepare resources still in use by another Pod. (#140212, @bart0sh) [SIG Node]
  • Fixed a race condition where server-side apply requests for custom resources could observe an updated CustomResourceDefinition before the apply path was fully synchronized, causing inconsistent dry-run behavior. (#140232, @googs1025) [SIG API Machinery]
  • Fixed a theoretical issue where nodes might have been denied access to synthesized ResourceClaims for pods using extended resources (for example, nvidia.com/gpu), causing containers to get stuck in ContainerCreating. Not observed in practice. (#138792, @dims) [SIG Auth and Node]
  • Fixed server-side apply to correctly drop status changes when tracking field ownership for PodGroup, PodCompositeGroup, and PodCertificateRequest. (#140654, @jpbetz) [SIG API Machinery, Auth, Scheduling and Testing]
  • Improved kubelet Topology Manager error messages when the prefer-closest-numa-nodes policy option is enabled on Windows nodes that do not expose NUMA distance information, clarifying that the option is not supported on those nodes. (#139760, @zylxjtu) [SIG Node]
  • Improved memory usage of kube-proxy by dropping the .metadata.managedFields field, which kube-proxy does not require. (#140056, @adrianmoisey) [SIG Network]
  • Logged a warning in the kubelet if a static Pod defines an invalid priority or priorityClassName. (#136705, @sreeram-venkitesh) [SIG Node]
  • Promoted apiserver_watch_events_total and apiserver_watch_events_sizes to Beta. (#137116, @tico88612) [SIG API Machinery, Instrumentation and Testing]
  • Removed RelaxedDNSSearchValidation feature gate. (#139217, @adrianmoisey) [SIG Apps, Node and Testing]
  • Removed locked GA feature gates RetryGenerateName, BtreeWatchCache, OrderedNamespaceDeletion, StreamingCollectionEncodingToJSON, StreamingCollectionEncodingToProtobuf, APIServerTracing, ResilientWatchCacheInitialization, and ConsistentListFromCache. (#138907, @Jefftree) [SIG API Machinery, Apps, Etcd and Node]
  • Removed the --concurrent-service-syncs kube-controller-manager flag (no-op since v1.31). (#138002, @Jefftree) [SIG API Machinery]
  • Removed the KubeletMinVersion label from the DRA e2e test covering multiple ResourceClaims. (#138001, @rogowski-piotr) [SIG Node and Testing]
  • Removed the PreventStaticPodAPIReferences feature gate. Static Pods can no longer reference API resources, and this behavior can no longer be disabled. (#140226, @sreeram-venkitesh) [SIG Node]
  • Stopped using maps for single-endpoint Services in kube-proxy nftables mode, increasing the speed of programming nftables. (#140723, @adrianmoisey) [SIG Network]
  • Switched StorageVersionMigration to use merge patch instead of SSA. (#138874, @michaelasp) [SIG API Machinery, Apps and Auth]
  • The SidecarContainers feature gate, unconditionally enabled since v1.33, is removed. (#137755, @HirazawaUi) [SIG Apps, Node, Scheduling and Testing]
  • The kube-apiserver --enable-logs-handler flag, deprecated in v1.15, is no longer marked deprecated. It remains off by default. (#138915, @BenTheElder) [SIG API Machinery]
  • The deprecated Alpha metrics apiserver_cache_list_total, apiserver_cache_list_fetched_objects_total, and apiserver_cache_list_returned_objects_total are no longer exposed by default. Consumers should migrate to the unified apiserver_storage_list_* metrics with the storage="watchcache" label. (#139154, @yedou37) [SIG API Machinery and Instrumentation]
  • Removed the no-op DefaultWatchCacheSize field of k8s.io/apiserver/pkg/server/options.EtcdOptions. (#134151, @ialidzhikov) [SIG API Machinery]
  • Updated DRA so that the ResourceClaim controller creates ResourceClaims from ResourceClaimTemplates referenced by a Pod that is a member of a PodGroup only when the DRAWorkloadResourceClaims feature gate is enabled. This prevents creating a ResourceClaim for an individual Pod when it is intended to be created for the PodGroup. (#138363, @nojnhuh) [SIG API Machinery, Apps, Node, Scheduling and Testing]
  • Updated the etcd client library to v3.6.11. (#138747, @humblec) [SIG API Machinery, Auth, Cloud Provider, Node and Scheduling]
  • kube-controller-manager and kube-scheduler both expose dynamic_resource_allocation_resourceclaim_creates_total as a metric for the number of ResourceClaims created, replacing the differently named metrics in each component. The kube-controller-manager metric resource_claims was moved to the same dynamic_resource_allocation subsystem. (#138542, @pohly) [SIG Apps, Instrumentation, Node, Release, Scheduling and Testing]
  • kubeadm: Removed the v1beta3 API which was deprecated since v1.31. The v1.35 kubeadm binary can be used to migrate to v1beta4 by using the command kubeadm config migrate. Additionally, removed the PublicKeysECDSA kubeadm-specific feature gate which was only kept for backwards compatibility with v1beta3. The support for ECDSA keys was added as part of the v1beta4 field ClusterConfiguration.EncryptionAlgorithm. Added a placeholder v1 API that is a copy of v1beta4 and is flagged as experimental and cannot be used yet. (#136016, @neolit123) [SIG Cluster Lifecycle]
  • kubeadm: Updated the supported etcd version to v3.6.10 for supported control plane versions v1.34, v1.35, and v1.36. (#138392, @humblec) [SIG API Machinery, Cloud Provider, Cluster Lifecycle, Etcd and Testing]
  • kubeadm: Updated the supported etcd version to v3.6.11 for supported control plane versions v1.34, v1.35, and v1.36. (#138746, @humblec) [SIG API Machinery, Cloud Provider, Cluster Lifecycle, Etcd and Testing]

Dependencies

Added

Changed

Removed



Contributors, the CHANGELOG-1.37.md has been bootstrapped with v1.37.0 release notes and you may edit now as needed.



Published by your Kubernetes Release Managers.

Reply all
Reply to author
Forward
0 new messages