net/tls: Fix race in TLS device down flow
authorTariq Toukan <tariqt@nvidia.com>
Fri, 15 Jul 2022 08:42:16 +0000 (11:42 +0300)
committerDavid S. Miller <davem@davemloft.net>
Mon, 18 Jul 2022 10:33:07 +0000 (11:33 +0100)
commitf08d8c1bb97c48f24a82afaa2fd8c140f8d3da8b
treee1eb102889f37c8cd9fce124b2c891157eabf7dc
parent613b065ca32e90209024ec4a6bb5ca887ee70980
net/tls: Fix race in TLS device down flow

Socket destruction flow and tls_device_down function sync against each
other using tls_device_lock and the context refcount, to guarantee the
device resources are freed via tls_dev_del() by the end of
tls_device_down.

In the following unfortunate flow, this won't happen:
- refcount is decreased to zero in tls_device_sk_destruct.
- tls_device_down starts, skips the context as refcount is zero, going
  all the way until it flushes the gc work, and returns without freeing
  the device resources.
- only then, tls_device_queue_ctx_destruction is called, queues the gc
  work and frees the context's device resources.

Solve it by decreasing the refcount in the socket's destruction flow
under the tls_device_lock, for perfect synchronization.  This does not
slow down the common likely destructor flow, in which both the refcount
is decreased and the spinlock is acquired, anyway.

Fixes: e8f69799810c ("net/tls: Add generic NIC offload infrastructure")
Reviewed-by: Maxim Mikityanskiy <maximmi@nvidia.com>
Signed-off-by: Tariq Toukan <tariqt@nvidia.com>
Reviewed-by: Jakub Kicinski <kuba@kernel.org>
Signed-off-by: David S. Miller <davem@davemloft.net>
net/tls/tls_device.c