Skip to content

boards/qemu-intel64: add citest configuration - #20030

Open
raiden00pl wants to merge 12 commits into
apache:masterfrom
raiden00pl:nuttx-qemu-pr
Open

boards/qemu-intel64: add citest configuration#20030
raiden00pl wants to merge 12 commits into
apache:masterfrom
raiden00pl:nuttx-qemu-pr

Conversation

@raiden00pl

Copy link
Copy Markdown
Member

Summary

CI/NTFC testing config for intel64

Impact

CI tests for intel64

Testing

CI

xiaoxiang781216
xiaoxiang781216 previously approved these changes Sep 1, 2026
@github-actions github-actions Bot added Area: Documentation Improvements or additions to documentation Size: M The size of the change in this PR is medium Board: x86_64 labels Sep 1, 2026
@github-actions

github-actions Bot commented Sep 1, 2026

Copy link
Copy Markdown

MemBrowse Memory Report

arduino-mega2560

  • flash: .text -14 B (-0.0%, 65,090 B / 262,144 B, total: 25% used)

esp32-devkitc

  • ROM: .flash.text -32 B (-0.0%, 124,888 B / 4,194,272 B, total: 3% used)
  • irom0_0_seg: .flash.text -32 B (-0.0%, 89,092 B / 3,342,304 B, total: 3% used)

mirtoo

  • kseg0_progmem: .text -20 B (-0.0%, 67,736 B / 131,072 B, total: 52% used)

qemu-armv8a

  • Code: .text.uart_recvchars -4 B, .text.uart_xmitchars -4 B, .text.work_qcancel +20 B (+0.0%, 336,128 B)

qemu-intel64

  • Code: .text -16 B (-0.0%, 8,659,721 B)

rx65n-rsk2mb

  • ROM: .text -16 B (-0.0%, 86,864 B / 2,097,152 B, total: 4% used)

s698pm-dkit

  • Code: .text -16 B (-0.0%, 365,216 B)

stm32-nucleo-f103rb

  • flash: .text -20 B (-0.1%, 34,536 B / 131,072 B, total: 26% used)
    No memory changes detected for:
  • hifive1-revb

xiaoxiang781216
xiaoxiang781216 previously approved these changes Sep 2, 2026
@raiden00pl

Copy link
Copy Markdown
Member Author

CI container uses old QEMU version that doens't support x2APIC, so we need this PR to run it #20044

xiaoxiang781216
xiaoxiang781216 previously approved these changes Sep 2, 2026
@github-actions github-actions Bot added the Arch: x86_64 Issues related to the x86_64 architecture label Sep 2, 2026
@acassis

acassis commented Sep 4, 2026

Copy link
Copy Markdown
Contributor

@raiden00pl please take a look:

=========================== short test summary info ============================
FAILED ../../nuttx-ntfc/external/nuttx-testing/arch/os/integration/test_arch_os_integration.py::test_os - "Device crashed" detected, during: call
FAILED ../../nuttx-ntfc/external/nuttx-testing/ltp/test_ltp_interface_integration.py::test_ltp_integration[ltp_interfaces_aio_cancel_2_1] - assert <CmdStatus.TIMEOUT: -2> == 0
FAILED ../../nuttx-ntfc/external/nuttx-testing/ltp/test_ltp_interface_integration.py::test_ltp_integration[ltp_interfaces_aio_cancel_2_2] - assert <CmdStatus.TIMEOUT: -2> == 0
FAILED ../../nuttx-ntfc/external/nuttx-testing/ltp/test_ltp_interface_integration.py::test_ltp_integration[ltp_interfaces_aio_cancel_8_1] - assert <CmdStatus.TIMEOUT: -2> == 0
FAILED ../../nuttx-ntfc/external/nuttx-testing/ltp/test_ltp_interface_integration.py::test_ltp_integration[ltp_interfaces_aio_cancel_9_1] - assert <CmdStatus.TIMEOUT: -2> == 0
FAILED ../../nuttx-ntfc/external/nuttx-testing/ltp/test_ltp_interface_integration.py::test_ltp_integration[ltp_interfaces_aio_error_1_1] - assert <CmdStatus.TIMEOUT: -2> == 0
FAILED ../../nuttx-ntfc/external/nuttx-testing/ltp/test_ltp_interface_integration.py::test_ltp_integration[ltp_interfaces_aio_error_3_1] - "Device busy_loop" detected, during: call
FAILED ../../nuttx-ntfc/external/nuttx-testing/ltp/test_ltp_interface_integration.py::test_ltp_integration[ltp_interfaces_lio_listio_10_1] - "Device crashed" detected, during: call
FAILED ../../nuttx-ntfc/external/nuttx-testing/ltp/test_ltp_interface_integration.py::test_ltp_integration[ltp_interfaces_lio_listio_15_1] - assert <CmdStatus.TIMEOUT: -2> == 0
FAILED ../../nuttx-ntfc/external/nuttx-testing/ltp/test_ltp_interface_integration.py::test_ltp_integration[ltp_interfaces_lio_listio_18_1] - assert <CmdStatus.TIMEOUT: -2> == 0
FAILED ../../nuttx-ntfc/external/nuttx-testing/ltp/test_ltp_interface_integration.py::test_ltp_integration[ltp_interfaces_lio_listio_1_1] - assert <CmdStatus.TIMEOUT: -2> == 0
FAILED ../../nuttx-ntfc/external/nuttx-testing/ltp/test_ltp_interface_integration.py::test_ltp_integration[ltp_interfaces_lio_listio_2_1] - assert <CmdStatus.TIMEOUT: -2> == 0
FAILED ../../nuttx-ntfc/external/nuttx-testing/ltp/test_ltp_interface_integration.py::test_ltp_integration[ltp_interfaces_lio_listio_3_1] - "Device busy_loop" detected, during: call
FAILED ../../nuttx-ntfc/external/nuttx-testing/ltp/test_ltp_interface_integration.py::test_ltp_integration[ltp_interfaces_lio_listio_7_1] - "Device crashed" detected, during: call
FAILED ../../nuttx-ntfc/external/nuttx-testing/ltp/test_ltp_interface_integration.py::test_ltp_integration[ltp_interfaces_pthread_attr_setdetachstate_2_1] - assert <CmdStatus.TIMEOUT: -2> == 0
FAILED ../../nuttx-ntfc/external/nuttx-testing/ltp/test_ltp_interface_integration.py::test_ltp_integration[ltp_interfaces_pthread_barrierattr_init_2_1] - assert <CmdStatus.TIMEOUT: -2> == 0
FAILED ../../nuttx-ntfc/external/nuttx-testing/ltp/test_ltp_interface_integration.py::test_ltp_integration[ltp_interfaces_pthread_barrierattr_setpshared_1_1] - assert <CmdStatus.TIMEOUT: -2> == 0
FAILED ../../nuttx-ntfc/external/nuttx-testing/ltp/test_ltp_interface_integration.py::test_ltp_integration[ltp_interfaces_pthread_barrierattr_setpshared_2_1] - assert <CmdStatus.TIMEOUT: -2> == 0
FAILED ../../nuttx-ntfc/external/nuttx-testing/ltp/test_ltp_interface_integration.py::test_ltp_integration[ltp_interfaces_pthread_cancel_4_1] - assert <CmdStatus.TIMEOUT: -2> == 0
FAILED ../../nuttx-ntfc/external/nuttx-testing/ltp/test_ltp_interface_integration.py::test_ltp_integration[ltp_interfaces_pthread_cancel_5_1] - "Device busy_loop" detected, during: call
FAILED ../../nuttx-ntfc/external/nuttx-testing/mm/cachetest/test_mm_cachetest_integration.py::test_cachetest_integration - assert <CmdStatus.TIMEOUT: -2> == 0
=========== 21 failed, 892 passed, 323 skipped in 3115.70s (0:51:55) ===========

At the end of the log there is also a Warning about CPython and gcc support to qemu

xiaoxiang781216
xiaoxiang781216 previously approved these changes Sep 7, 2026
@github-actions github-actions Bot added Area: Drivers Drivers issues and removed Area: Documentation Improvements or additions to documentation labels Sep 7, 2026
/* Send while we still have data in the TX buffer & room in the fifo.
*
* uart_putxmitchar() advances xmit.head from thread context without
* holding the critical section, so on SMP it can move (and wrap) while

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

but the critical section is held at line 63

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

the producer writes head outside any lock. This ring is a single-producer/single-consumer ring:

if (nexthead != dev->xmit.tail)
{
/* No.. not full. Add the character to the TX buffer and return. */
dev->xmit.buffer[dev->xmit.head] = ch;
dev->xmit.head = nexthead;
break;
}

Comment thread libs/libc/aio/lio_listio.c
Comment thread sched/wqueue/kwork_cancel.c Outdated
@raiden00pl
raiden00pl force-pushed the nuttx-qemu-pr branch 4 times, most recently from 95727ab to c6dd1a3 Compare September 8, 2026 14:09
CI/NTFC testing config for intel64

Signed-off-by: raiden00pl <raiden00@railab.me>
intel64_oneshot_start() takes g_oneshot_spin and then, if the timer is
already running, calls intel64_oneshot_cancel(), which takes the same
spinlock again.  Spinlocks are not recursive, so the CPU spins forever
on its own lock while holding the critical section; the HPET timer ISR
on another CPU then blocks on g_cpu_irqlock and the system hangs.

This is hit as soon as the tickless scheduler re-arms a running HPET
oneshot timer under SMP (ostest task_restart, LTP aio tests).

Stop the running timer inline instead of calling cancel: disable the
interrupt, detach the ISR so up_enable_irq() does not assert on a busy
IRQ, and clear the running flag.  The ISR, comparator and interrupt
enable are reprogrammed by the rest of the function anyway.

Assisted-by: Claude Code
Signed-off-by: raiden00pl <raiden00@railab.me>
work_cancel() used to return -ENOENT when the work structure was not
in the queue, and callers depend on that: aio_cancel() tears down the
AIO container (file_put() + aioc_free()) only when work_cancel()
reports success, because a work item that is not queued may already be
executing on a worker thread (see the comment in fs/aio/aio_cancel.c).

Since commit 6f72f54 ("sched/wqueue: Refactor delayed and periodical
workqueue") work_cancel() returns OK unconditionally, and commit
d2e01b9 ("sched/wqueue: harden custom queue lifecycle") kept that
behaviour and dropped -ENOENT from the function documentation.  Under
SMP the LTP aio_cancel tests then free the aio container and its file
while the lpwork thread is still executing aio_write_worker() on it,
which ends in a page fault in file_write() (f_inode == NULL) and a
panic.

Return -ENOENT again when the work is not queued, and document it.
For the synchronous variant "not queued" alone does not tell whether
the callback is running: the worker scan does, so report OK when a
running callback was found and waited for, and -ENOENT only when the
work was neither queued nor running.

Assisted-by: Claude Code
Signed-off-by: raiden00pl <raiden00@railab.me>
fix nxstyle issues in intel64_hpet.c

Assisted-by: Claude Code
Signed-off-by: raiden00pl <raiden00@railab.me>
intel64_hpet_setisr() with a NULL handler detached the ISR with
irq_attach(irq, NULL), which installs irq_unexpected_isr().  The oneshot
driver does this every time the timer expires or is re-armed, so an HPET
interrupt already in flight to another CPU lands on the unexpected ISR
and panics the system:

  irq_unexpected_isr: ERROR irq: 34

seen under SMP with the LTP test suite.  Just mask the interrupt and keep
the ISR attached; intel64_oneshot_handler() already treats an interrupt
that arrives while the timer is not running as spurious.

Assisted-by: Claude Code
Signed-off-by: raiden00pl <raiden00@railab.me>
fix nxstyle issues in lio_listio.c

Assisted-by: Claude Code
Signed-off-by: raiden00pl <raiden00@railab.me>
lio_listio(LIO_NOWAIT) submits the requests first and only afterwards
installs the SIGPOLL handler and attaches the per-request private data
in lio_sigsetup().  Under SMP the requests complete on the low priority
work queue while this is going on, and aio_signal() queues SIGPOLL for
each of them as it finishes:

- A signal delivered for a request that lio_sigsetup() found already
  finished (aio_priv left NULL) made lio_sighandler() dereference a
  NULL private pointer (LTP lio_listio 10-1, 15-1, 2-1 crash with a
  page fault in lio_sighandler on qemu-intel64 SMP).
- A signal delivered between the aio_result check and the aio_priv
  store of the last outstanding request was consumed without private
  data, no later signal arrived, and the caller was never notified.
- If every request had completed by the time the loop ran, no handler
  invocation could ever see the private data and the caller hung.
- After the completing handler had notified the caller, signals still
  queued for other requests ran the handler again with their own copy
  of the private data.

Block SIGPOLL while the handler is installed and the private data is
attached, so that the completion signals are delivered only once every
entry is set up.  Ignore a signal for a request without private data,
notify the caller directly from lio_sigsetup() when everything had
already completed, and detach the private data of every entry once the
list is done so that late signals find nothing to do.

Assisted-by: Claude Code
Signed-off-by: raiden00pl <raiden00@railab.me>
The "cancel everything on this descriptor" loop in aio_cancel() only
moved to the next container inside the branch taken when work_cancel()
succeeded.  When work_cancel() reports -ENOENT because a worker thread
is already executing the request, the loop restarted its search from
the same container and spun forever, hanging the task (and, since the
scan runs with interrupts enabled but never yields, that CPU) without
any output.  This was hit by the LTP aio_cancel tests on SMP.

Capture the next container before calling work_cancel() so the loop
always makes progress.

Assisted-by: Claude Code
Signed-off-by: raiden00pl <raiden00@railab.me>
uart_putxmitchar() advances xmit.head from thread context without
holding the critical section, so on SMP the head index can move, and
wrap around, while uart_xmitchars() runs in the TX interrupt on another
CPU.  Since commit b319c27 ("serial: Added APIs for receiving and
sending multiple chars") the sendbuf path of uart_xmitchars() reads
xmit.head twice: once to decide whether the pending data is contiguous
and again to compute its length.  If the producer wraps the index in
between, the computed length goes negative, is passed to sendbuf() as a
huge size_t and the driver transmits memory far beyond the ring buffer.

The per-byte path reads the index only once and is not affected, which
is why this went unnoticed: the batch path is only used by drivers that
implement sendbuf, and the 16550 driver gained it in commit 45c38d8
("drivers/serial/16550: add polling mode support for serial drivers").
qemu-intel64 with SMP is the first configuration combining a sendbuf
driver with a producer running on another CPU.

On qemu-intel64 SMP this shows up as an endless stream of NUL bytes on
the console (captured with gdb: head = 1, tail = 8, size = 16, and
u16550_sendbuf() called with size = (size_t)-7), which makes the ntfc
test harness fail every test that runs while the flood lasts.

Read the head index once per loop iteration and use that snapshot for
both the contiguity test and the length.  The producer only ever moves
the index forward, so a stale snapshot merely sends less now.

Assisted-by: Claude Code
Signed-off-by: raiden00pl <raiden00@railab.me>
Same issue as the previous commit, on the receive side: uart_read()
advances recv.tail from thread context without holding the critical
section, but the recvbuf batch path of uart_recvchars() reads recv.tail
several times (the full check, the watermark count and the free-space
computation).  If uart_read() moves and wraps the index in between, the
computed free space goes negative and is passed to recvbuf() as a huge
size_t, which lets the driver store past the end of the ring buffer.

Read recv.tail once per loop iteration and derive everything from that
snapshot.  The consumer only ever moves the index forward, so a stale
snapshot merely stores less now.

Assisted-by: Claude Code
Signed-off-by: raiden00pl <raiden00@railab.me>
intel64_oneshot_handler() cleared oneshot->handler and oneshot->arg
after picking them up, without holding g_oneshot_spin, while
intel64_oneshot_start() re-arms the timer under that lock from another
CPU.  Now that the HPET ISR stays attached across a re-arm, a stale
interrupt can interleave with start(): it reads the freshly installed
handler, clears it, and start() then sets running = true again, so the
genuine expiry that follows finds running == true with a NULL handler
and jumps to address zero from interrupt context (page fault at RIP 0
in the CPU0 IDLE task while the LTP lio_listio tests were running), or
the alarm is simply lost and the tickless system stops.

The handler and its argument are owned by start() and cancel(); the ISR
only needs to read them.  Leave them alone in the ISR and skip the call
if none is installed.  The remaining effect of a stale interrupt is an
early invocation of the alarm callback, which is harmless: the tickless
scheduler re-evaluates its expirations and re-arms the timer.

Assisted-by: Claude Code
Signed-off-by: raiden00pl <raiden00@railab.me>
The aio worker functions started by decanting the container, which
frees the container back to the pool and drops the file reference,
and only then performed the I/O through the aiocb.  Nothing was
holding the file open while the read/write/fsync was in progress, and
the request was no longer on the pending list, so a concurrent close()
or aio_cancel() could pull the file from under the running operation.

Keep the container until the operation has completed and decant it
just before signalling completion.

Assisted-by: Claude Code
Signed-off-by: raiden00pl <raiden00@railab.me>
@github-actions github-actions Bot added Size: L The size of the change in this PR is large and removed Size: M The size of the change in this PR is medium labels Sep 8, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Arch: x86_64 Issues related to the x86_64 architecture Area: Drivers Drivers issues Board: x86_64 Size: L The size of the change in this PR is large

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants