| CVE |
Vendors |
Products |
Updated |
CVSS v3.1 |
| In the Linux kernel, the following vulnerability has been resolved:
mm, swap: fix SWAP_USAGE_OFFLIST_BIT collision with real usage count
SWAP_USAGE_OFFLIST_BIT is embedded in the si->inuse_pages usage counter,
and is meant to sit above any value that counter can reach. However, it
is defined from BITS_PER_TYPE(atomic_t), so it is bit 30. On a system
with 4 KiB pages the flag collides with the usage count once that count
reaches 4 TiB.
swap_usage_in_pages() masks bit 30 out, so whenever the real count has
that bit set, every caller of it reads 4 TiB low:
* /proc/swaps understates Used by 4 TiB.
* A raw count of exactly 2^30 masks to zero, so try_to_unuse() takes its
"if (!swap_usage_in_pages(si)) goto success;" early exit and swapoff
tears the device down while pages are still swapped out. Nothing in
the rest of swapoff aborts the teardown, so those pages are lost.
Independently of swapoff, the collision also corrupts the counter and the
plist. On a device in normal use, a free that leaves bit 30 set in the
count makes swap_usage_sub() see the flag where there is only count, and
call add_to_avail_list(). It clears the bit with
fetch_and(~SWAP_USAGE_OFFLIST_BIT), leaving the stored count 4 TiB below
the real one, and calls plist_add() on a device that is already listed,
tripping the WARN_ON(!plist_node_empty(node)) in plist_add() and linking
the node a second time.
Change the definition of SWAP_USAGE_OFFLIST_BIT to be based on
atomic_long_t instead. Note that the usage counter field itself is of
this same type, so it is still a valid bit. |
| In the Linux kernel, the following vulnerability has been resolved:
mm/shrinker: fix bogus set_shrinker_bit() with cgroup.memory=nokmem
With cgroup.memory=nokmem, shrinker_memcg_alloc() bails out early and
never allocates an id, so shrinker->id keeps the 0 it got from the
kzalloc() in shrinker_alloc(). __list_lru_init() then copies that 0 into
lru->shrinker_id, where it looks like a valid bit index.
Nothing calls expand_shrinker_info() on nokmem either, so shrinker_nr_max
stays 0 and every memcg ends up with an empty map (map_nr_max == 0).
deferred_split_folio() hands a real memcg to __list_lru_add() regardless
of whether the lru is memcg aware, so the first THP queued in a cgroup
does set_shrinker_bit(memcg, nid, 0) and trips the bounds check:
WARNING: mm/shrinker.c:212 at set_shrinker_bit+0x7d/0x90, CPU#126
Call Trace:
<TASK>
deferred_split_folio+0x18c/0x220
map_anon_folio_pmd_nopf+0xdd/0x130
map_anon_folio_pmd_pf+0x14/0xb0
do_huge_pmd_anonymous_page+0x1a1/0x620
__handle_mm_fault+0xea9/0x10d0
handle_mm_fault+0xe5/0x320
do_user_addr_fault+0x1cc/0x870
exc_page_fault+0x81/0x1b0
asm_exc_page_fault+0x27/0x30
</TASK>
Harmless, the WARN_ON_ONCE() is what keeps the out of bounds unit[] read
from happening, but the id should not look valid in the first place.
Clear it before returning.
Two other spots could paper over this: drop the id in __list_lru_init()
when nokmem turns memcg_aware off, or make deferred_split_folio() pass
NULL like list_lru_add_obj() does. Both leave shrinker->id lying around
for the next caller, so fix it where the id is handed out. |
| In the Linux kernel, the following vulnerability has been resolved:
mm/vma: correctly unaccount on mmap_prepare() failure
__mmap_setup() accounts memory for relevant mappings via:
security_vm_enough_memory_mm()
-> __vm_enough_memory()
-> vm_acct_memory()
If __mmap_setup() fails, this indicates that this accounting did not take
place, and thus it's appropriate for __mmap_region() to jump to
abort_munmap.
However if call_mmap_prepare() fails, it also jumps there and any accounted
memory is not correctly unaccounted.
Fix this by handling each error separately. |
| In the Linux kernel, the following vulnerability has been resolved:
mm: filemap: retain mapped dropbehind folios
Fault-around can map ready dropbehind folios without going through the
normal page-cache lookup that clears dropbehind. A mapping represents a
competing cached user, so retain the folio instead of forcibly unmapping
it when writeback completes.
For a mapped folio, folio_unmap_invalidate() can call
unmap_mapping_folio(), which takes i_mmap_rwsem and may sleep. Retaining
mapped folios avoids this path when folio_end_dropbehind() runs in
non-preemptible task context.
Tal was able to trigger a sleeping-in-atomic warning due to this [1].
Unmapped dropbehind folios continue through the existing invalidation path. |
| In the Linux kernel, the following vulnerability has been resolved:
KEYS: encrypted: fix integer overflow of datablob_len
encrypted_key_alloc() stores datablob_len in a u16. It is computed from
multiple string and payload lengths. If the result exceeds U16_MAX, the
assignment truncates the allocation size. KASAN reports a 32760-byte
slab-out-of-bounds write when __ekey_init() copies the master key
description into the undersized buffer.
The total payload length stored in key->datalen is also a u16. Use
check_add_overflow() to reject values that do not fit either destination,
and use kzalloc_flex() for the flexible-array allocation. |
| In the Linux kernel, the following vulnerability has been resolved:
KEYS: trusted: Fix tpm2_load_cmd() boundary check
tpm2_load_cmd() does boundary checks against the ASN.1 size i.e.,
payload->blob_len. Address this by passing the decoded blob size to
tpm2_load_cmd(), and use it for the boundary checks. |
| In the Linux kernel, the following vulnerability has been resolved:
sched_ext: Fix NULL sched deref in kfunc sub-sched error paths
When the root scheduler has sub-scheds attached, the COMPAT kfunc
wrappers scx_bpf_select_cpu_and() and scx_bpf_dsq_insert_vtime() refuse
the call and report to @p's scheduler:
scx_error(scx_task_sched(p), "... must be used");
The wrappers are reachable with tasks that have no scheduler.
scx_bpf_select_cpu_and() is in the select_cpu kfunc group, which
scx_kfunc_context_filter() opens to BPF_PROG_TYPE_SYSCALL programs;
scx_bpf_dsq_insert_vtime() is in the enqueue_dispatch group, which
ops.enqueue() and ops.dispatch() may call with any KF_RCU task -- the
group has no kf_tasks validation, and scx_dsq_insert_preamble() checks
task ownership with scx_task_on_sched() precisely because @p may be an
arbitrary task.
scx_task_sched(p) is p->scx.sched, which is NULL for tasks past
sched_ext_dead() -- which clears it via scx_disable_and_exit_task() on
exit -- and for idle tasks, which the enable paths skip as they are
never scheduled through SCX. It is also an rcu_dereference_protected()
that expects @p's pi_lock or rq lock, which neither wrapper holds.
Passing NULL to scx_error() reaches scx_vexit(), which dereferences
sch->exit_info, oopsing the kernel.
One concrete trigger exercised while developing the fix: a
BPF_PROG_TYPE_SYSCALL program calling the select_cpu_and wrapper on an
exited-but-not-reaped task while a sub-scheduler was attached (its pid
stays findable while the zombie is unreaped; faulting instruction is
the scx_vexit() prologue "mov r15,[rdi+0x398]" with RDI=NULL and 0x398
the offset of sch->exit_info):
sched_ext: BPF scheduler "kfunc_subsched_null" enabled
sched_ext: BPF sub-scheduler "kfunc_subsched_null" enabled
sched_ext: Unassociated program run_select_cpu_ (id 76)
BUG: kernel NULL pointer dereference, address: 0000000000000398
#PF: supervisor read access in kernel mode
#PF: error_code(0x0000) - not-present page
Oops: Oops: 0000 [#1] SMP NOPTI
CPU: 7 UID: 0 PID: 8201 Comm: kfunc_test_runn Tainted: G W
RIP: 0010:scx_vexit+0x25/0xa0
Code: ... <4c> 8b bf 98 03 00 00 ...
CR2: 0000000000000398
Call Trace:
<TASK>
__scx_exit+0x4f/0x70
scx_bpf_select_cpu_and+0xab/0xb0
bpf_prog_430ed61a7b66e03a_run_select_cpu_and+0x9c/0xe7
? __x64_sys_bpf+0x2c/0x40
bpf_prog_test_run_syscall+0x130/0x2f0
__sys_bpf+0x930/0x10d0
? __x64_sys_bpf+0x2c/0x40
__x64_sys_bpf+0x2c/0x40
do_syscall_64+0xbc/0x460
entry_SYSCALL_64_after_hwframe+0x76/0x7e
</TASK>
Read @p's scheduler under RCU instead, which the wrappers can do from
their guard(rcu)(): fault it when it can be determined, and when it
can't be determined -- @p is a task past sched_ext_dead() or an idle
task -- there is nothing obviously wrong to report, so just refuse the
call as before without faulting any scheduler.
These COMPAT wrappers are scheduled for eventual removal once the
deprecation grace period elapses, but until then -- and regardless of
their removal timeline -- they must not oops the kernel on a task they
are handed. |
| In the Linux kernel, the following vulnerability has been resolved:
sched_ext: Close the pre-enable ops error claim window
scx_alloc_and_add_sched() publishes ops->priv before
scx_root_enable_workfn() switches the state to SCX_ENABLING. An error
claimed via scx_bpf_error_bstr() from an associated BPF program in that
window is consumed by scx_disable_workfn(), which takes the pre-enable
shortcut in scx_root_disable(). The shortcut returns without any teardown
and restores SCX_DISABLED with an unconditional scx_set_enable_state() xchg
racing the enable workfn's own transition. The enable then completes with
the claim consumed: the scheduler stays up but can never be disabled again,
and bpf_scx_unreg() frees it while still in use, resulting in a
use-after-free. Both WARN_ON_ONCE()s fire back to back:
WARNING: kernel/sched/ext/ext.c:7522 at
scx_root_enable_workfn+0xeec/0x1be0, CPU#3: scx_enable_help/276
WARNING: kernel/sched/ext/ext.c:6398 at scx_root_disable+0xb50/0xdb8,
CPU#0: sched_ext_helpe/664
scx_root_enable_workfn() switches to SCX_ENABLING before the scheduler
allocation, so ops->priv is never visible while SCX_DISABLED. The allocation
failure path restores SCX_DISABLED. |
| In the Linux kernel, the following vulnerability has been resolved:
i2c: atr: fix dangling adapter pointer on add failure
i2c_atr_add_adapter() stores atr->adapter[chan_id] before
i2c_add_adapter() so that the I2C bus notifier can match child clients
during registration. On failure the channel is freed but the slot was
left pointing at freed memory, which can lead to use-after-free in
i2c_atr_del_adapter() / cleanup and also block reuse with -EEXIST.
Clear the slot on the i2c_add_adapter() error path before freeing chan. |
| In the Linux kernel, the following vulnerability has been resolved:
IB/mlx4: Fix use-after-free on pkey sysfs registration failure
register_pkey_tree() ignores errors from register_one_pkey_tree() and
continues registering the remaining slaves. The per-slave error path has
already released the pkey parent kobjects, but their pointers remain
stored in the device. A later device cleanup therefore passes the stale
pointers to kobject_put(), causing a use-after-free.
Clear the parent pointers after releasing a failed slave tree and skip
unregistered trees during device cleanup. This preserves the existing
best-effort registration behavior while preventing a second cleanup of
the failed tree. |
| In the Linux kernel, the following vulnerability has been resolved:
IB/hfi1: Fix the PIO_CRED credit-return mmap
hfi1_file_mmap()'s PIO_CRED case must hand user space the single
credit-return page that holds this context's entry. That page is the
second or third page of the per-node credit-return allocation once the
hardware send context index reaches 64 or 128, so the failure below is
intermittent: when the entry lands on the first page the offset is zero
and everything works.
Two things are wrong.
First, cr_page_offset is a byte offset but .va is a struct
credit_return *, so adding it is pointer arithmetic and scales the offset
by sizeof(struct credit_return) == 64. memvirt then lands 256 KiB or
512 KiB past a 10240-byte allocation. With an IOMMU translating, that
address is inside the vmalloc range but in no vm_area, so
dma_mmap_coherent() -> iommu_dma_mmap() finds no pages, vmalloc_to_pfn()
returns page_to_pfn(NULL), and remap_pfn_range() installs a frame above
MAXPHYADDR. The first user read then takes:
psm2_ep_open_pr: Corrupted page table at address 7a14d007e000
PGD 800000013886a067 P4D 800000013886a067 PUD 13886b067 PMD 13886c067
PTE 800049168e911235
Oops: Bad pagetable: 000d [#1] SMP PTI
Second, and still wrong once the arithmetic is corrected,
dma_mmap_coherent() describes a whole coherent buffer and selects the
page within it with vma->vm_pgoff. Offsetting cpu_addr has no effect:
for a vmap'd allocation iommu_dma_mmap() uses cpu_addr only to locate the
vm_area and then maps pages[vm_pgoff], which hfi1_file_mmap() has just
set to 0. User space therefore always receives the first credit-return
page, every credit read is for the wrong context, and send PIO stalls
forever.
Use the DMA API as intended: pass the base of the allocation with its
full length and select the page with vm_pgoff. A separate length is
needed because memlen must keep describing the VMA for the existing size
check. The dma-direct path stays correct as well, since dma_direct_mmap()
adds the same vm_pgoff to the base pfn.
Tested on a Dell T7610 (Xeon E5-2650 v2, Intel IOMMU in DMA-FQ mode)
against a Threadripper PRO 3995WX peer, both Omni-Path 100. Before this
change psm2_ep_open() Oopses the kernel; with only the arithmetic
corrected psm2_ep_open() succeeds but any transfer that uses send PIO
hangs, PSM2_SDMA=2 (send PIO disabled) completing normally while
PSM2_SDMA=0 (send PIO only) hangs every time. With this change send PIO,
send DMA and the default mixed mode all work. |
| In the Linux kernel, the following vulnerability has been resolved:
selinux: preserve user SID across nested backing files
SELinux saves the user file SID in a backing-file security blob so it
remains available after mmap() replaces vma->vm_file with a backing file.
For nested backing files (overlayfs over overlayfs, or FUSE passthrough
backed by overlayfs), user_file may itself be a backing file. Its
fsec->sid is the SID of the mounter that opened it, rather than the user
that opened the top-level file. mprotect() then checks fd { use } against
the mounter SID. This can incorrectly deny access without a domain
transition, or check the wrong target SID after one.
Copy the saved user SID when user_file is a backing file. Keep using the
regular file SID for the first backing layer.
With two nested overlayfs mounts and SELinux enforcing,
mprotect(PROT_READ) returns EACCES with an fd { use } denial against the
mounter SID. With this change, mprotect() succeeds.
Tested on arm64 QEMU with a small BusyBox initramfs and a purpose-built
SELinux policy. The original test was also repeated with Fedora Cloud
Base 44 userspace and gave the same result. |
| In the Linux kernel, the following vulnerability has been resolved:
selinux: recheck intermediate backing files on mprotect()
mprotect() can be used to bypass the SELinux checks that mmap() performs
against the intermediate layers of a stacked filesystem.
mmap() checks every backing layer as the request descends through the
stack. mprotect() only has the lowest backing file in vma->vm_file, so it
rechecks the top-level user and the lowest mounter, but skips the mounters
of every layer in between. With two nested overlayfs mounts and a policy
denying mounter_t -> middle_file_t:file { execute }, a direct
mmap(PROT_EXEC) is denied:
avc: denied { execute } for pid=71 comm="nested_exec"
path="/payload" dev="overlay" ino=9
scontext=user_u:base_r:mounter_t
tcontext=user_u:object_r:middle_file_t tclass=file permissive=0
while mmap(PROT_NONE) followed by mprotect(PROT_EXEC) succeeds.
Preserve each intermediate path, mounter SID and file-description SID in
the backing-file security blob, copying the saved entries when another
backing layer is opened. Allocate the array only for nested backing files,
and release it and the path references in the backing_file_free hook.
During mprotect(), recheck fd { use } and the requested inode permissions
for every saved mounter, and include the intermediate layers in the execmod
checks. Policy for nested stacking may then need to grant intermediate
mounters what a direct mmap() already requires, and execmod on intermediate
labels for binaries using text relocations.
Tested on arm64 QEMU with a small BusyBox initramfs and a purpose-built
SELinux policy, on a mainline tree containing
commit f2381b546e7e ("fs: fix user path of nested backing files").
[PM: subject tweak] |
| In the Linux kernel, the following vulnerability has been resolved:
mmc: core: Cancel SDIO IRQ work before freeing host
A host controller that uses sdio_signal_irq() schedules host->sdio_irq_work
from its interrupt handler. That work is only cancelled on the suspend
path (mmc_sdio_suspend()), not on the remove/free path, so a worker armed
just before the controller freed its IRQ can run after
mmc_host_classdev_release() has freed the host and dereference it through
container_of().
Cancel host->sdio_irq_work in mmc_free_host(), like the existing
host->detect drain added by commit 1036f69e2513 ("mmc: core: Cancel
delayed work before releasing host").
This issue was found by an in-house static analysis tool. |
| In the Linux kernel, the following vulnerability has been resolved:
mmc: hsq: Fix use-after-free in retry work
mmc_hsq_pump_requests() queues retry_work when request_atomic() returns
-EBUSY; today sdhci-sprd is the only consumer that implements
request_atomic(). The work is embedded in a devm-allocated mmc_hsq, but
is never cancelled during driver removal. Work still pending at unbind
can therefore run after the devm allocation has been released and
dereference hsq->mmc and hsq->mrq.
Use devm_work_autocancel() to cancel and drain retry_work before the devm
allocation is released. By the time devres cleanup begins,
mmc_remove_host() has already stopped the host, so no new requests can
arm the work.
This issue was found by an in-house static analysis tool. |
| In the Linux kernel, the following vulnerability has been resolved:
mmc: mmci: Fix use-after-free in busy-timeout work
ux500_busy_complete() can queue ux500_busy_timeout_work for an R1b
command, but mmci_remove() never cancels it. The work can subsequently
dereference the devm-allocated mmci_host after it has been released.
Mask the controller interrupts and disable the delayed work during
removal. This drains any queued instance and stops an IRQ handler that
is still in progress from queueing the work again once it has been
disabled.
This issue was found by an in-house static analysis tool. |
| In the Linux kernel, the following vulnerability has been resolved:
mmc: mxcmmc: cancel data work and watchdog on remove
mxcmci_remove() frees the host through the devm tail, but neither it nor
mmc_remove_host() drains the driver's own asynchronous state.
host->watchdog, a 10 s timer armed on the DMA path in mxcmci_setup_data(),
is deleted only by the DMA- and IRQ-complete paths, which the remove path
does not explicitly drain; it can therefore fire after the host is freed
and dereference it in mxcmci_watchdog(). host->datawork, armed from the
IRQ handler on the PIO path, is not cancelled by the remove path either.
Free the devm-registered IRQ, then cancel datawork and delete the watchdog
in mxcmci_remove(), before dma_release_channel(). Freeing the IRQ first
keeps a trailing handler from re-arming datawork between the cancel and
the host free. Both callbacks are non-self-rearming.
This issue was found by an in-house static analysis tool. |
| In the Linux kernel, the following vulnerability has been resolved:
mmc: sdhci-of-aspeed: Remove children before releasing SDC resources
Probe failure and removal leave SDHCI child devices registered after the
parent clock and managed resources are released.
Unregister the OF children in reverse order before disabling the parent
clock on both paths. Use of_platform_device_destroy() because manual
child creation does not set the flag required by of_platform_depopulate().
This issue was identified during our ongoing static-analysis research
while reviewing kernel code. |
| In the Linux kernel, the following vulnerability has been resolved:
mmc: sdio_uart: fix xmit_fifo leak when the port table is full
sdio_uart_add_port() allocates the transmit fifo before claiming a
slot in sdio_uart_table[]. When all UART_NR slots are taken, it
returns -EBUSY with the fifo still allocated, but the probe error
path only kfree()s the port, leaking the transmit fifo.
Free the fifo in the failure path of sdio_uart_add_port() itself so
the function retains nothing on error. |
| In the Linux kernel, the following vulnerability has been resolved:
mmc: spi: reset bytes_xfered before retrying CRC failures
mmc_spi_data_do() updates data->bytes_xfered after each block has been
transferred successfully. If a later block in the same data request
fails with a CRC error, data->bytes_xfered may therefore contain the
number of bytes completed before the failing block.
mmc_spi_request() has a private recovery path for such CRC failures. It
sends STOP_TRANSMISSION, clears data->error and jumps back to
crc_recover to issue the same command and data request again. However,
it does not clear data->bytes_xfered before the retry.
If the retry succeeds, the request is completed with the bytes from the
failed attempt still included in data->bytes_xfered. For a multi-block
request this can make the completed request report more bytes than were
transferred by the successful retry, and can even exceed the request size
when most blocks completed before the CRC error.
This is most likely to be observed on MMC-over-SPI systems where long
multi-block transfers occasionally hit a data CRC error but the
mmc_spi-internal retry succeeds. The data itself is retried, but the
completion accounting is not.
Clear data->bytes_xfered together with data->error before repeating the
request so the final completion reports only the bytes transferred by the
successful attempt. |