1: 38a8e1d9d0d1 ! 1: baec8a1a72ec wifi: ath11k: release peer accounting on peer delete timeout @@ Metadata ## Commit message ## wifi: ath11k: release peer accounting on peer delete timeout  - On some deployments access points intermittently stop accepting new - station associations after several hours of uptime with frequent - roaming/reconnects. The kernel logs - - ath11k_pci ....: failed to create peer due to insufficient peer entry resource in firmware - - and hostapd reports "Could not add STA to kernel driver". A "wifi - down/up" (radio restart) on the affected radio restores service. - - Despite the message text, this is not a firmware peer-table exhaustion. - The message is emitted by the driver-side gate in ath11k_peer_create(): - - if (ar->num_peers > (ar->max_num_peers - 1)) - - i.e. the driver's own ar->num_peers accounting has leaked and reached - the ceiling. It is always preceded by a peer-delete that timed out: - - ath11k_pci ....: invalid vdev id in peer delete resp ev 1 - ath11k_pci ....: Timeout in receiving peer delete response - ath11k_pci ....: failed to delete peer vdev_id .. addr .. ret -110 - ath11k_peer_delete() only decrements ar->num_peers when - __ath11k_peer_delete() returns 0. On a delete timeout + __ath11k_peer_delete() returns 0. On a peer-delete-confirmation timeout __ath11k_peer_delete() returns -ETIMEDOUT, so the decrement is skipped - and one num_peers slot is leaked per event. After max_num_peers such - timeouts ath11k_peer_create() rejects every new station until the radio - is restarted. - - In the observed case the peer-unmap event is received normally (there is - no "failed wait for peer deleted" log), so ath11k_peer_unmap_event() has - already removed the peer from ab->peers and freed it; only the num_peers - counter is left wrong. The delete-response completion is missed because - the response event is dropped in ath11k_peer_delete_resp_event() when - ath11k_mac_get_ar_by_vdev_id() cannot resolve the vdev ("invalid vdev id - in peer delete resp ev"), so ar->peer_delete_done is never signalled and - the second wait in ath11k_wait_for_peer_delete_done() times out. - - Return success from __ath11k_peer_delete() on the timeout path so that - ath11k_peer_delete() releases the num_peers slot. As a safety net also - drop the local peer if it is still on the list; that only happens in the - other timeout case, where ath11k_wait_for_peer_deleted() itself timed out - and no unmap event removed the peer. The peer has already been removed - from the rhash earlier in __ath11k_peer_delete(), so only the list - removal and free remain, mirroring ath11k_peer_unmap_event(). A late - unmap event would then simply fail to find the peer id and log a harmless - warning instead of touching freed memory. - - The problem was reproduced deterministically with an out-of-tree debug - patch that adds module parameters to force - ath11k_wait_for_peer_delete_done() to return -ETIMEDOUT and to cap - max_num_peers: after max_num_peers such deletes the AP permanently - rejects new stations, and with this change it keeps accepting them. + and one ar->num_peers slot is leaked. Once enough slots have leaked the + ar->num_peers > (ar->max_num_peers - 1) gate in ath11k_peer_create() + rejects every new station with "insufficient peer entry resource in + firmware", and the AP stops accepting associations until the radio is + restarted with a wifi down/up. + + This was observed in the field on an AP after hours of uptime with + frequent roaming and reconnects: clients could no longer associate even + though the firmware peer table was not actually exhausted, only the + host-side ar->num_peers accounting had leaked. + + The delete-confirmation timeout itself is otherwise harmless. The host + first waits for the peer unmap event in ath11k_wait_for_peer_deleted(), + so by the time the delete-response completion times out the peer has + already been removed from ab->peers and freed by + ath11k_peer_unmap_event(). Only the ar->num_peers counter is left + inconsistent. + + The timeout is reached when the peer unmap event arrives but the peer + delete response is missed, e.g. because the response event is dropped in + ath11k_peer_delete_resp_event() on an unresolved vdev id (logged as + "invalid vdev id in peer delete resp ev"). Do not treat the delete + confirmation timeout as fatal so that ath11k_peer_delete() releases the + ar->num_peers slot instead of leaking it.  Fixes: 690ace20ff79 ("ath11k: peer delete synchronization with firmware") Cc: stable@vger.kernel.org @@ Commit message  ## drivers/net/wireless/ath/ath11k/peer.c ##  @@ drivers/net/wireless/ath/ath11k/peer.c: static int __ath11k_peer_delete(struct ath11k *ar, u32 vdev_id, const u8 *addr) + return ret; }  - ret = ath11k_wait_for_peer_delete_done(ar, vdev_id, addr); +- ret = ath11k_wait_for_peer_delete_done(ar, vdev_id, addr);  - if (ret)  - return ret; -+ if (ret) { -+ /* -+ * The delete timed out. The peer is normally already freed by -+ * the unmap event; drop it here only if it is still on the -+ * list. Either way return success so that ath11k_peer_delete() -+ * releases the num_peers slot instead of leaking it. -+ */ -+ spin_lock_bh(&ab->base_lock); -+ peer = ath11k_peer_find(ab, vdev_id, addr); -+ if (peer) { -+ list_del(&peer->list); -+ kfree(peer); -+ } -+ spin_unlock_bh(&ab->base_lock); -+ } ++ /* Ignore the return value: the peer is already freed, only its ++ * ar->num_peers slot would otherwise leak on a delete timeout. ++ */ ++ ath11k_wait_for_peer_delete_done(ar, vdev_id, addr);  return 0; }