#uACPI - a portable and easy-to-integrate ACPI implementation
1 messages · Page 41 of 1
its a result of the limine template globbing (and uacpi being inside the src dir)
qemu-system-x86_64 -cdrom orange.iso -M q35 -enable-kvm -m 1G
i dont getting errors when compiling
wtf
I also get basically the same score
interesting
did u commit the allocator fixes
try this iso
with this arguments
now i need to think
i got a much better result
why we getting another results
i see an iso so i ran it
like why they getting 60k but you and me getting 2mln
what CPU do you have?
i have a ryzen 5800x
ryzen 5 3600
code unoptimized for intel
intel discrimination
git submodule skill issue
more like glob skill issue lol
for now just delete tests folder ✅
why not just move it out of the src folder and either add the uacpi sources manually or do a different glob with only c files in the uacpi dir
maybe i can just remove a tests folder in makefile
although this is also a normal solution
could someone (with a fast cpu) test this iso 🥺
qemu-system-x86_64 -M q35 -enable-kvm -cpu host -debugcon stdio
and send me the score afterward
@left orbit test iso pls 🥺
on it
ty
[uACPI][INFO]: successfully loaded 1 AML blob, 1705 ops in 1ms (avg 1144648/s)
[uACPI][INFO]: successfully loaded 1 AML blob, 1705 ops in 1ms (avg 1108243/s)
[uACPI][INFO]: successfully loaded 1 AML blob, 1705 ops in 1ms (avg 1132732/s)
[uACPI][INFO]: successfully loaded 1 AML blob, 1705 ops in 1ms (avg 1119295/s)```
seems kinda low
2167942/s is what I got
yeah wsl moment
can you send screenshot?
it varies a little but its around there
what a fast cpu can do to a kernel
@fiery turtle I got a new score for you
for obos
[uACPI][INFO]: successfully loaded 1 AML blob, 1705 ops in 1ms (avg 2141179/s)```
which cpu
14600k?
i5 13600k
ah not ev en close lol
lol
@mortal yoke would you mind testing this image?
qwinci do I add your kernel?
slowly but surely I will take over the leaderboard
if so, do i literally call it crescent 2
~6.1M but it varies like streaks said too
yeah, that is kinda expected tbh
I have to improve my debugging tools so I can profile and get a good score on AMD CPUs too
i need to fix nixpkgs nooooo
idk lol, the description that I posted is kinda weird at least (and the name is what it is but ig it doesn't really matter even if it would be called crescent2 for now and be renamed later™️)
(maybe start by actually implementing some debugging tools)
the description seems fine
sure np
i didnt see that u edited it as a full entry
infy is uacpi_kernel_install_interrupt_handler only called during namespace init, and not in load
idr
@fiery turtle if you do update the scores this is my submission for now until I can coerce someone with an even better cpu to run it
its only called in uacpi_initialize
or actually looks like namespace load too 
i think load only(?)
it should be calld during event init
yeah I accidentally looked at the wrong function lol
what would happen if "notable projects using uacpi" was changed to actually only include notable projects
like ones that actually use uacpi API, and don't just have it for the lulz
lol
I'd bet half the kernels on there would be gone 
some of these might get abandoned and then ill remove them
like astral 
astral is kinda back
I don't use uacpi api 💀
didn't leo said he wouldn't eat or sleep and try to get more points if managarm was surpassed 
so you just have it for the lulz
though I do intend on using it after acpi.sys becames a thing
to implement acpi ioctl api
that was back when 1M was considered insane

good ol' times
I should make a time machine using obos
then go back in time
and tell old me how to optimize stuff to the brim
in obos
it'll crash mid-travel and you'll end up with dinosaurs
shi you're right
use NT acpi.sys maybe
for funs
use the uacpi acpi.sys from ros (once it becomes upstream) 
could be fun, would have to figure out what kind of interface it has tho
I still think everyone should test on the same HW and I still offer my services to do so.
what do you have?
we should have an osdev performance test server
I used to run a lot of the benchmarks on my 14700K, not a new idea
like how browsers have wpt.fun or whatever it was (web platform tests)
we are all in favour of establishing a clearinghouse for benchmarks
the question is who will get it set up
I don't have something that can be up all the time
if you do it manually that's also acceptable i suppose
just a bit of an onerous demand on someone
Manually is fine by me as long as it's not like 10 a day
I want to run hydra for my OS (to build the entirety of nixpkgs)
then I can give a % of working nix pkgs
Should we make a poll of whoms't becomes the tester™️?
an auto runner is not an absurd idea it's just that realistically whoever owns it will want to be careful who he lets use it, because it'd be a lot more work to make it robust for public exposure
take a look at yesterdays discussion
OSDev test lab
im totally down but it shouldnt be a willy nilly "ill test lol"
I think it's fine for like tagged releases
my idea was to have it running as a systemd service so basically whenever my pc is on i can test in the background lmao
I actually have a computer that can probably be on like 24/7
and host something like that
although it's not fast
it has like a pentium 4 right
if it's useful I have a pi5 for native aarch64
lmfao
i think we don't need like an actual 24/7 thing tho
we dont
that'll be nice
i'll only be competing with managarm there
we just need to accumulate a few isos and test once in a while
i have an m1 mac i don't use 
it would take half an hour tops
I use(d) it for mc servers
but a pi would be nice
fair enough, as many people i am also volunteering
i can reboot into linux for host consistency every few days
first person to suggest and implement an actual infra for that gets selected 
I also call this pc test subject 2 iirc
so we want it to be automatic after all? 🤔
like ideally it's:
- an automatic submission, at least for projects already on the board
- automatic batch PR for score changes
also would be nice to have isolcpus= and pin qemu there
but that code isn't upstream
Inb4 GitHub hooks that submit the new ISO to score 
you can do a pr using a github user token so
yeah
also automatic testing would require getting the score from smth like serial and also ideally doing an average because there are some very bad runs sometimes
checkout master, make changes, commit and push to remote then open pr
like 30 lines of python code
yeah we talked about that
yep
hyperfine it
ok now it is upstream
and of course pin on one cpu
I meant submitting to the server that scores them
basically require anyone to provide it over debugcon or serial
ye
and probably a timeout of like 10s
ok let me play around with it lol
uTournament
uci is a chess engine thing
if anyone actually makes a proper thing we can put it on the uacpi project page as well for nice versioning
i meant like
uEgoBooster
i meant to reply to this
that's kinda what it is lol
ultraContIntegragitiion
also imagine a separate application page for new projects to submit themselves
hacker@raptor:~/uacpi-bench-runner$ ls /mnt/c/Users/hacker/Desktop/ | grep iso
Apartheid.Linux.Cyberwar.Edition.x64.iso
archlinux-2025.02.01-x86_64.iso
astral.iso
barebones (1).iso
barebones.iso
davix (1).iso
davix.iso
davix.iso.gz
debian-live-12.9.0-amd64-standard.iso
image-x86_64.iso
iso9660.iso
iso9660.isoasdads
iso9660.iso.zip
obos.iso
proxima (1).iso
proxima.iso
testbed.iso
Win11_24H2_English_x64.iso```
which one of these isos is most likely to do uacpi output over serial
Ignore the first one
😇
ugh I want to do TCP but at the same time I don't want to
I want to because it's fucking cool
obos will do ⭐
proxima should
I feel like this should be mostly manual to avoid malicious ISOs
ok let me try proxima obos was kinda slow
I don't want to because it fucking sucks
recent proxima builds do debugcon
the review process is manual ofc
yep my proxima iso does debugcon
and existing users can auth via github to re-submit or smth
epiccc
have you ever heard of ls *.iso
the classic <any command that can probably grep itself> | grep
cat foo.txt | grep bar
cat /dev/nvme0n1p1 | grep 
any idea how i can capture both serial and debugcon to a file?
or two separate files, doesnt matter
idk how this whole chardev shit works
-debugcon file:shit.txt
one to stdio, one to file:/dev/stdout
output to separate files
yep
you will be reading it from python anyway right
shell script 
python ftw
u can also just open pipes via subprocess
so yeah what I usually do is write all my programs in C because I don't know any sort of scripting language properly
but idk how to make qemu output serial to a pipe sooo
or debugcon for that matter
astral does debugcon
yeah dont do that anyway, there are some memes with those pipes overflowing
Unix socket?
obos automatically enables debugcon if it detects it's on a hypervisor
BadgerOS does both debugcon and COM1
wouldn't qemu_system_x86_64 -serial stdio | command work
wait until u find out about in(0xE9) == 0xE9
theres also pipe:some_pipe_name output for qemu
whoaaaaaaaaaaaaaaaaaaaaaaaa
very anti pythonic way to capture output lmfao
works in both bochs and qemu
you can still do -serial stdio and use the subprocess.run thing to pipe no?
ye
i know for windows you can go pipe:\\.\pipe\<pipename> id assume theres similar for using a unix socket on *nix or something idk
named pipes?
yeah
yeah posix has those
(dont use that pipe bit for windbg btw it makes qemu die)
yeah i only know the windows ones lol
gl
im in oh god why does irl keep giving me things i need to do i just want to do nvme mode lol
for my os
i think i might have to do unix sockets because i can just read from them from python and i can exit once i get the measurement
idk if i can just open the files and read until i get the measurements too
I will be implementing it based on the first TCP spec 
hopefully TCP is backwards compatibile
i need to write some async python code for this shit i think, i need the timeout but i also need to read lines from both of the files
idk if it's the first, but I'm doing it on RFC793
@fiery turtle ur the python expert here, any idea how to do this properly?
i kinda hate subprocess piping hmm
well im just thinking of how to do the timeout + reading both files part
i will kill the process either after getting the measurement or after timeout
why not let the submission tell you which output it uses?
why do u not like the idea of just
subprocess.run(qemu -debugcon <file1> -serial <file2>)
with open(<file1>) as f:
e9_output = f.read()
with open(<file2>) as f:
serial_output = f.read()
because i dont know when the output comes in
so i might not read anything
i have to wait for the data to come in
but i also want a timeout
subprocess run is synchrnonous
i know
u can set e.g. a 10 second timeout and just wait until it dies
run takes in a timeout right
but i cant do anything until the timeout expires
and i dont want to do that
i want to exit as soon as the measurements are output by the kernel
it'd be funny if a ten second timeout isn't enough for some kernels
i am gonna do some async memes and see if i can make it work
cuz of how slow they are
lfg
say how you do plan on implementing this
ns load used to take like 30 seconds on nyaux?
if output_measured: == True:
exit()

if you want to check the output as it comes in for early exit you should probably use Popen instead of run
we dont run random shit tho, only verified projects should be able to auto submit
fair
I feel like assuming that submitted ISOs aren't cheating is fine
how do you know I won't 
decompile the binary
[uACPI][INFO]: successfully loaded 1 AML blob, 1705 ops in 0ms (avg 154523545/s)
[uACPI][INFO]: namespace initialization done in 1ms: 36 devices, 0 thermal zones```
if the score is sus
obos fastest kernel
we'll make it so the runner injects random aml salt that will be printed if your run is fair 
it's easy to implement too
Level1Techs has a SiFive P550 Premiere that they want to give remote access to apparently
if we make automated speedtest
we can probably ask them
we could just inject an SSDT that invokes an undefined reference, which will be printed by uACPI in your kernel's stdout
the name of the ref will be printed
and it doenst incur any unfair score penalties either
also make the ref randomly generated for good measure
wouldn't that include logging performance in the benchmark?
yeah thats what im saying
hm ig
but kernel_log is also an API soo 
actually, no, if we make it an _INI method, then it will only get to run after the benchmark
if it gets run we'll also know the number of times u did init and if u did namespace_initialize at all
wait I just realized
well yeah but unlike mutexes and allocators performance varies wildly based simply on where the kernel logs to - serial is slower than debugcon, framebuffer might have to scroll, etc
it's been 1 year since the first port of uacpi to a kernel
yeye scratch that
ah ok
the ssdt can be literally
// no cost, one extra namespace node for all runs
Method (\_INI) {
Debug = \RAND.OM.NAME.<HASH>
}
I just realized that uacpi didn't have an overridable stdlib.h header when I first ported it
lol for some reason select returns readable all the time on the file handles
i think i need a unix socket or something
Can't you send some sort of signal to qemu?
then this code can be injected via -acpitable <file>
To send power button event
why would i
i dont care about shutting down the vm
i want to wait for measurements asynchronously
while also having a timeout
we could require that for a "complete" impl tbh lol
OS shutting down would be also a good way to know that the impl is working 
yeah true
but then not all os's may want to shut down immediately on power button press
true
No, but like wouldn't it arrive once uACPI is loaded?
or have that implemented yet even with a lot of other acpi stuff implemented
yeah
L kernel
or propagates them to userspace
astral propagates it to userspace
propegates lol
obos' kernel literally ignores all events other than EC events and wake GPEs
and wake GPEs don't even get handled by the kernel
i dont really care about power button stuff atm for my kernel
focusing on getting it working and doing stuff while on lol
reading through old messages of me testing uacpi in its early days is always fun
when infy was surprised it could actually shutdown on real hw
real
any idea how i can properly do the measurement part? like filter out any outliers and such
and then average out the good ones
yeah but like, you can still get outliers
yeah yeah i mean
it can be done by a wrapper script
easy enough to switch out qemu itself
for a wrapped version
true
now you test it on other people's shitty kernels 
But at least, out of 16, at least 1 should be working on any given hardware
maybe initializing uACPI, version x
is better
now the question is, should this be printed out even for early table init, or only during uacpi_initialize
imo, call it on the first call of uacpi initialize or of early table init
i can make like a PRINT_ONCE for it and stuff but still
You can print the version next to the benchmark
it might also help to print the uacpi in uacpi errors after uacpi init
for bug reports or smth
obos does that on panic
i mean if it's printed on boot anyway
u can just paste the entire log
hmm yeah looks better at the very top
uOEMID lol
ok i got the first version of this thing, basically running the iso 50 times and collecting the measurements from serial and debugcon lol
mean = statistics.mean(measurements)
stdev = statistics.stdev(measurements)
lower_bound = mean - stdev * 2
upper_bound = mean + stdev * 2
filtered = [x for x in measurements if lower_bound <= x <= upper_bound]
mean = statistics.mean(filtered)```
it also does this shit for some reason idk how this works 
and it writes the statistics to a json file for further processing i guess
damn u even got component decoupling
honestly for printing the version, I would do something like uACPI version 1.0.1
idk what that means
but yes! (maybe, if u say so)
you can git clone it and enable debugcon by creating user/config.h and adding #define CONFIG_DEBUGCON 1
effort
yea lol
now time for the cpu pinning memes
I should probably like, actually parse a kernel commandline and provide the option to enable both debugcon and serial through it
but effort 
@fiery turtle you said cgroups is the way?
well no
run qemu with debugthreads=on or whatever, grep for the one with nameVCPU/0, pin it to the allocated core with taskset or similar
thats the simplest way
ideally u also pin the main thread there, or allocate a separate core for it
well you don't want to run the benchmark script on the same cpu
ideally qemu should get majority of the cpu time there
-name debug-threads=on i think?
no no
like
when u start linux, offline a CPU, or offline it via a kernel command line via isolcpus
oh what the fuck
then pin qemu there so it gets literally the entire cpu to itself
thats the fairest way
honestly i dont want to go too overboard with this whole thing
considering it's gonna be running on my personal system right
true ig
yes it is a good idea to make it as fair as possible
but i don't want to get into weird isolation stuff like that
u can still pin it to some cpu, it will just be random whether it gets a lot of time
not really that weird but yeha
it's not weird, it just excludes it from the scheduler list of cpus to schedule on
well it should, assuming there isn't much going on on the system
only if there was a way to do it at runtime...
hmm
that will halt it i think
yeah i guess
aka u wont be able to schedule on it
after using isolcpus, you can apparently just use taskset to "pin" it
yeah thats what im saying
thats nice
it's like a guarded area
but yeah im not finding any info about isolcpus at runtime
cpuset is apparently also something worth looking into
you can also make qemu pause on startup and give it infinite timeslices (aka SCHED_FIFO with highest priority) before unpausing it
is that enough to guarantee maximum cpu time though?
processes with sched_fifo do not have timeslices, and if they're also in the highest priority they don't get preempted
so as soon as qemu gets a timeslice it'll keep running until it blocks for whatever reason
so, now, how do i do the "SCHED_FIFO" thing and highest priority thing? 
though with this especially you'll want to be careful to not run it on the same cpu as the runner script, otherwise the timeout may break
sched_setscheduler
also yeah pin your script on some other cpu lol
whatever cpu affinity you give qemu, give yourself the inverse of it
Since i'm making the register api public i've decided to add some comments, hopefully this is as clear as possible
any idea how i can unpause qemu without any input..?
cont
monitor
ideally make it spawn a qmp socket and talk over that
true
ideally uacpi should have tests like that
with a real qemu integration for some kernel, testing shutdown, pci hotplug, hotunplug, resource parsing, etc
for that qmp events are a must
also added this notice lol
okay i got the scheduling stuff done, now i need to figure out how to properly set the affinity for qemu and the python process
idk which cpu it should be pinned to tho lol
0?
isn't that like the worst cpu you can choose though
usually most interrupts are directed there
maybe on a hobby kernel
at least on wsl cpu 0 has the most irqs
how did u check?
I would avoid picking the BSP, which is usually (if not always) CPU0.
also, I doubt that isolcpus= would allow you to isolate the BSP
i am not doing isolcpus, but yeah bsp is not the best choice
i think the cpu to pin is pretty machine specific
indeed
if u have /sys/kernel/debug/u should have irq stuff there, which can show u how many vectors are allocated on each cpu and stuff
if you are running WSL, the interrupt stuff probably isn't very meaningful.
like, "it's a real Linux kernel" etc... but it doesn't actually interact with the real physical hardware
yeah, no
In my experience (and I've done a lot of benchmarking in my career), isolcpus and pinning doesn't make a difference for single threaded workloads, provided that you don't run anything else concurrently that consumes significant amounts of cpu time
isolcpus mostly makes a difference when you want to have low latency when reacting to a device interrupt
yeah this feels like hyperoptimizing something that doesn't need that much thought put into it, i think SCHED_FIFO + nice -20 will be good enough
just to make sure it gets the most cpu time as possible
yup, probably.
yeah especially if u run it many times
And pinning does make a difference for parallel applications where you have many threads doing computations at the same time
I don't think nice values make a difference for SCHED_FIFO
how does the Linux scheduler's task migrator cope with a FIFO task eating 100% CPU time? will the FIFO task move around between CPUs, or will the other tasks be kicked off from the CPU? something to consider
(for our purposes, we would prefer the latter)
i think that's how SCHED_FIFO works
I don't think nice affects whether you get preempted or not
SCHED_FIFO: First in-first out scheduling
SCHED_FIFO can be used only with static priorities higher than 0,
which means that when a SCHED_FIFO thread becomes runnable, it
will always immediately preempt any currently running SCHED_OTHER,
SCHED_BATCH, or SCHED_IDLE thread. SCHED_FIFO is a simple
scheduling algorithm without time slicing. For threads scheduled
under the SCHED_FIFO policy, the following rules apply:
[...]
No other events will move a thread scheduled under the SCHED_FIFO
policy in the wait list of runnable threads with equal static
priority.
A SCHED_FIFO thread runs until either it is blocked by an I/O
request, it is preempted by a higher priority thread, or it calls
sched_yield(2).```
Also, sched_fifo processes can't get preempted
it definitely does
nice might affect if/how you preempt others when unblocking
Nice affects how often a thread is scheduled, not how long it is scheduled
read last sentence
idk
yes it doesn't have a time slice but it can still get preempted by a higher priority thread
idk how that works i am basically braindead but yeah
Yes but there are no higher priority threads if you give it the highest priority
i do that with renice... so yeah
nice -20 is the highest priority
SCHED_FIFO on it's own does not change the priority
right?
as i said, idfk how that works lmao
Nice is not priority
hm okay
Priority is sched_priority in sched_param
I personally wouldn't bother with sched_fifo either
At least not without extensive testing
As it's way easier to make a mistake than it is to do it right
For example, if qemu's IO threads run with SCHED_FIFO, this may have a significant detrimental effect on perf
proxima actually runs a bit faster without SCHED_FIFO
interesting
2.3-2.4 -> 2.5M
at least with a random iso i had on my desktop lol
wait wtf why is it so slow? doesn't it have the record at like 8M or something?
lol astral has that too
lol the iterative namespace walk api from ACPICA that Haiku uses doesn't check if a node is temporary
they forgor to add a check 💀
pmos doesn't print uacpi logs to serial
not surprising considering linux doesnt use it
WAAAAAAAAAAAAAAAA
Give me 10 mins
git clone https://github.com/dbstream/davix.git
cd davix
mkdir user
echo '#define CONFIG_DEBUGCON 1' > user/config.h
echo '#define CONFIG_UACPI 1' >> user/config.h
./run --limine
assuming I didn't typo anything (I typed this on a phone) and that any reasonable gcc and binutils is accessible as x86_64-elf-*, this should just work™️, and it will generate an ISO and place it in /tmp/bootable.iso
tbh I should change my buildsystem to autodetect gcc/binutils and use the system gcc/binutils as a fallback (printing a warning ofc)
day 0 without finding a new crazy bug in ACPICA
❌ any reasonable gcc and binutils is accessible as x86_64-elf-*
anyway i reported it to the haiku people
you can just ln -s /usr:bin/gcc /usr/local/bin/x86_64-elf-gcc and do the same for ld
that's what I use on my machine for developing lol
but effort

and I will probably unbork it tomorrow by actually, like, parsing the kernel command line
i absolutely am
WS
'd
fuck arithmetic mean
was it a math skill issue?
there are multiple sample points >11M but the 6-7M samples just kill them
it is a WSL moment
unfortunately
wtf
it would be interesting to have, in addition to the mean, like 10/25/75/90th percentile
if you provide me with the math, sure 
what did you do?
as i have already noted on multiple occasions: I am Brian dead
I have no idea
- store results in a list.
- sort the list.
that's literally it
well i already do that
because this is usually what happens when you remove a directory that is the cwd of a bash instance
bash does weird CWD things so eg. navigating upwards (..) from a symlink works as if there was no symlink
Is ns16550 interrupt level triggered?
(I'm just gonna hardcode it for now...)
Although I guess ACPI should provide that
dumb question
I think that weird things can also happen if you rename a directory that is part of the path to CWD of a bash instance
for the Nth percentile, take (N * (num_measurements-1)) / 100 and use that to index into the list.
{
"raw_measurements": [
6094335, 6682788, 6647536, 10104182, 6362320, 6744675, 6697042, 6834269,
6716035, 6684596, 6671675, 6658829, 6124589, 10084878, 6827346, 6770144,
6579531, 6770440, 6673503, 6790907, 6672510, 5999169, 10136802, 9916999,
7033683, 10015154, 6815883, 6711515, 5204374, 10137224, 6557616, 10108915,
10094730, 6221946, 6750496, 6744088, 6668622, 6289841, 6736707, 10111073,
10059353, 6231565, 6947189, 4575105, 6642356, 6265111, 6745609, 6702281,
6714686, 6115824
],
"filtered_measurements": [
6094335, 6682788, 6647536, 10104182, 6362320, 6744675, 6697042, 6834269,
6716035, 6684596, 6671675, 6658829, 6124589, 10084878, 6827346, 6770144,
6579531, 6770440, 6673503, 6790907, 6672510, 5999169, 10136802, 9916999,
7033683, 10015154, 6815883, 6711515, 5204374, 10137224, 6557616, 10108915,
10094730, 6221946, 6750496, 6744088, 6668622, 6289841, 6736707, 10111073,
10059353, 6231565, 6947189, 4575105, 6642356, 6265111, 6745609, 6702281,
6714686, 6115824
],
"lower_bound": 4237969.691588384,
"upper_bound": 10219872.148411617,
"mean": 7228920.92,
"stdev": 1495475.614205808,
"p75": 6862499.0,
"p90": 10103236.8,
"p95": 10122651.05,
"p97": 10137000.34,
"p99": 10137430.78
}
simd proxima
if num_measurements is 101, this works out quite nicely
i misread 11M, wasn't actually there so that's mb
u can see how bad the variance is on WSL, damn
4.2M-10.2M
Will it handle a bunch of text correctly?
Like if it just printf the full initialization sequence
i mean yeah, why not
it just looks for the specific log from uacpi
UACPI_AVG_OPS_REGEX = re.compile(
r"successfully loaded \d+ AML blob, (\d+) ops in (\d+)ms \(avg (\d+)/s\)"
)```
no need for all the groups but yeah
it prints blobs if there was more than one 
trolled regex
on my AMD CPU (my kernel has AMD skill issues), I usually get like ~4.6M, but I sometimes get the occasional 3M. so that's understandable.
but here, there are lots of ~6M, ~7M, ~10M, a 5M and a 4M.
- r"successfully loaded \d+ AML blob, \d+ ops in \d+ms \(avg (\d+)/s\)"
+ r"successfully loaded \d+ AML blobs?, \d+ ops in \d+ms \(avg (\d+)/s\)"
fixed
but i dont expect multiple blobs with q35
ok there's some weird shit going on with the math that i dont feel like figuring out at the moment cuz im tired as fuck
but somehow python is hallucinating numbers
statistics.quantiles is returning a value that shouldn't be there, to be precise 7100247.6 for the 99th percentile while non of the values go above 7M lol
statistics is a built in python module btw ^
maybe it is because there is the 98.Xth percentile and the 99.Yth percentile but no exact 99th percentile, and thus the prev and next are combined according to some weight
but none of the values went above 7M so idk where it got that from
but i will investigate after an 8h nap
perhaps it is actually 'inventing' new values by doing some stddev shit
Bruh, my RISC-V driver was good enough to work as is on x86 (after I added port io instead of mmio) 💀
anyway, build as it is in dev branch
@left orbit
*took me 40 
(this driver is asynchronous and interrupt driven for both reads and writes
)
oh we're testing?
run it 100 times and take the median
if you have a problem in statistics that you don't know how to solve (because you know nothing about statistics just like me), median is the answer
ns16550 is still ns16550 eh?
pick whatever is most convenient for you
its your os? pick the best score
its someone else's? pick the worst score
@fiery turtle mint told me you use your own memcpy instead of the system one? and that you need to override it? what is that about
how am I supposed to do that from Ada
you need to override it for good performance
in your kernel api header, you need to #define uacpi_memcpy __builtin_memcpy
(or can, at least)
^ repeat this for memset, memcmp, memmove and strlen
(or the subset that you have, anyway)
if you dont want to you dont have to
but its a pretty significant perf diff afaik
Why does uacpi do that? It seems stupid
so that the OS doesn't have to provide them if they don't have them i think?
the others make more sense tbh
like printf
GCC can and will generate calls to memcpy even in freestanding code
Same for clang
this is actively harmful for any non C project, we are now to have C headers alongside our own ABI implementations? that is so unwieldy
its a 5 line header
Yes, so kernels have to provide memcpy anyway
and its optional
Specifically for the memcpy/memset/memmove it's a bit redundant tbh
Given the builtin versions already exists
but yeah ^^
memmove isn't really needed (as in calls to it aren't generated by the compilers as far as I know)
its something I need to wire in my Ada build system
god knows how
that's why i did not do it
ah, makes sense
i just didn't because i did not wanna bother lol
it was gonna be a pita so i left it as default
especially for no reason
tbh you can just add -Duacpi_memcpy=__builtin_memcpy to the CFLAGS i think?
the default is not even that bad, like the score stays about the same
There are also some str* functions I think
I mean, the default is basically a simple mem* impl
And I kinda doubt tbh that there are any huge copies being done in the code
i propose that uacpi switches to this implementation of memcpy: c void memcpy(void* dst, void* src, unsigned long long size) { typedef struct { char c[size]; } fuckingwhat; *(fuckingwhat*)dst = *(fuckingwhat*)src; }
Actually, can't the functions just be made weak?
I just dont understand why this is not part of kernel provided ABI
no, that's unportable
also i suspect it messes with LTO
i think a lot of the performance gain is not from overriding them, but from gcc/clang knowing its memcpy?
just ignore it its not strictly speaking necessary
why would you not handle this yourself internally then with some ifdefs
(clang doesnt support this wonderful extension btw)
as a compiler-specific optimization
It's an optional abi, like sized frees
why are optional abis done in such an uncomfortable way for non C bindings
tbh all stuff that needs ifdefs or headers is going to be annoying to deal with in Ironclad's build system
be it this or anything else
does ironclad's build system not support adding an include directory?
Because no one came and said it was uncomfortable outside of c binding most likely
or a -D cflag
yes
it can be done
i just don't quite understand either why this is the way it is, personally, but whatever
Overriding the header with a user header is meant as a way to more easily allow to override the functions instead of needing to have in the buildsystem -Duacpi_memcpy=..., and maybe some impls would need other kernel headers which would then need to be included as well
like what i mean is that i don't understand what's gained by uacpi shipping its own versions of these functions by default
instead of hard relying on them
it's easier if you dont care about the performance gain?
I think just for easier build
Oh wait
They also support msvc
Idk if msvc has something like __builtin_memcpy that magically works
Even tho that is probably fixed by just either using the builtin or using a custom prototype that the compiler expects
And still requiring the kernel to impl it
i think this design choice comes from the fact that it used to contain sprintf?
the point is that you have a prescribed kernel glue layer of like 25 functions already, adding 5 more that everyone pretty much already has sounds like a no brainer to me
true
Like specifically for the memcpy/memset/memcmp/memmove functions I think using the normal functions could work as is
There are some string functions that depends on the kernel might not exist
you still need to include SOME header for memcpy/memset/etc though?
You can just define them to be the expected versions directly
you cant though
Like have uacpi define the prototype
okay how do you write the prototype?
You can, because the compiler expects these functions and these signatures
on 64bit sysv, you either write memcpy as void* memcpy(void* dst, const void* src, unsigned long size); or as void* memcpy(void* dst, const void* src, unsigned long long size);
you use size_t
Literally as void* memcpy(void*,void*,size_t)
That's what the compiler defines it to be
does uacpi include stddef?
ah
Maybe
lol
yes if you don't override it
it includes freestanding headers yes
it has its own overridable version that delegates to stddef.h by default
Even if not there are compiler builtins for the type
and yes i would depend on the standard prototypes already, not like uacpi_ prefixed ones
if i use UACPI_OVERRIDE_LIBC, where would i need to put that header?
Which uacpi already checks for the specific compiler
they include "uacpi_libc.h" iirc
so. at the include path
Now as for the vsnprintf, that might be much different
but they are not necessarily the same as the type used in memcpy, aiui
yeah
i wouldn't depend on host for vsnprintf
for several other reasons
especially because its (partially) exposed to aml
Yeah it has its own vsnprintf, but you can override it the same way
mainly because it's not perf sensitive where it would even matter
so it needs to be exactly accurate for compat
yeah proxima for example uses non standard printf format specs
memcpy is pretty much C ABI with how compilers work nowadays
These are the overridable functions
defining it yourself is dumb when you are meant to be embedded
its just asking for messes
just assume they are present, and use them, done
extern memcpy(whatever)
solved the issue
Worst case for GCC/clang use the builtin for msvc use whatever msvc specifically wants
that's what flanterm does
that's what i am proposing
and flanterm does well
that is the easiest, most powerful, least pain in the ass thing you can do
everything else is just bad design and working around things outside the real world
point me to an implementation hosting uacpi that doesnt have memcpy
there isn't gonna be one
uacpi is asking for allocators, for interrupt and event handling, and for memory mapping, but asking for memcpy is too much
and if you wanna do the __builtin_memcpy trick btw you can still do it wholly inside uACPI with ifdefs
for compilers that support it
at which point the dependency is implicit
note that in some cases __builtin_memcpy instead of plain memcpy degrades performance
well idk, that's something for infy to figure out when to use lmao
I think if you mark the real mem* as non-inline it solved the problem
It just requires you to use the __builtin everywhere
on -O3 LTO clang with memcpy implemented in C, using the prototype instead of builtin increased proxima's score by like 1M iirc
I dont even ncessarily care about the performance, which I am sure I will get a massive boost of
this is just bad design
not a single time that the builtin uacpi memcpys are used its intended behaviour
not a single system that supports uacpi doesnt have a memcpy
its just working around ghost problems that dont exist and adding complexity for no reason
it's fine if he wants to keep the option but it should probably be opt-in, not out
imho
like if he wants to use it for unit tests or something
where it may make sense, idk
huh
strange because for me, replacing the builtin with a prototype makes like zero difference
are you using LTO?
GCC doesn't exhibit this behavior last time I tried
although my memcpy is literally just
- shuffle arguments
rep movsbret
no, just plain -O2
that would make much more sense
and if you're not using LTO and/or not implementing memcpy in C the __builtin is much better
then you wont exhibit this behavior anyway i think
like code inside uACPI could realistically cause a dependency on memcpy
like, directly, from the compiler
at that point i don't see why this whole circus is even a thing
iirc infy tries to make sure that doesn't happen
tries means nothing
just implement uacpi_memcpy like this: c void uacpi_memcpy(void* dst, void* src, unsigned long long size) { typedef struct { char noThisIsNotSupportedByClang[size]; } fuckingwhat; *(fuckingwhat*)dst = *(fuckingwhat*)src; }
wtf does that mean even
he can try now, but a future version may make it happen
LLVM can insert a memcpy wherever it wants
or a different architecture
and you cannot stop that
you can reduce it tho
not a valid argument
why though
yeah idk for memcpy specifically it doesn't make that much sense
you either have or not have a dependency, it's not a spectrum
reducing it while still depending on it means nothing
by using the compiler you always implicitly depend on it
just like how gcc and clang allow themselves to use libgcc/compiler-rt whenever they want, but it's still completely viable to write kernels in a way that doesn't use libgcc
Isn't rep movsb kinda bad for performance most of the time?
technically it isn't viable
its bad for small copies
its something doable tho
like you can avoid popcnt and whatever other things like that
linux does it so compilers are practically guaranteed to stay sane in their libgcc uses
Like I think just manual loop unrolling is fast(er)
like sure there's no real guarantee
idk what you are talking about
but
but no GCC or clang version is gonna get released that cannot build linux
on non-x86-64 it is absolutely not viable
and linux doesn't use libgcc
hey
let me try to express myself
one second
Linux does things like shipping its own freestanding headers so i am not surprised if it also ships its own libgcc compat routines
that's a different thing
uacpi literally calls libgcc though
for popcount
(this is something that the limine-c-template does too for example)
(with gcc at least)
linux doesn't provide libgcc, and uacpi doesn't use it anymore
are you sure linux doesn't provide libgcc routines at all?
yep
idkf, maybe it is.
will try it out later today
they have a separate macro for 64 bit division f.ex
well that sounds like them taking the stupid route
you cannot, however, stop gcc from calling its own 64 bit division
something like cc-runtime is painless to use
Linux could use it totally fine
or compiler-rt
and when I worked on integrating uacpi into linux for a bit I encountered the popcount stuff
like the phd memcpy is the way it is for a reason 
you can, as evidenced by the fact that linux does it
eh there isn't a huge diff between it and rep movsb
no, what they do is they TRY
with builtin memcpy falling back to that for larger copies
but relying on that is unportable, version dependent, and can break on any update
you can't get a documented guarantee that it'll never be called by a future compiler, but you can get a practical guarantee because no future compiler will get released in a state where it cannot compile linux
also this isnt a strong guarantee
it doesn't matter lol
linux could be implemented in a slightly different way
that would make it not hit it
and you would
not depending on mem* is not sane on gcc/clang
also that
just like not depending on libgcc or alternatives is
oh yeah mem* is a different story
i mean... https://godbolt.org/z/34sY86v1r
yeah
you cant not depend on memcpy
and memset
you also cant really not depend on the rest
That was removed a very long time ago
fucking what?
that's a legal struct definition in C?
yes
its not supported by clang though
typedef struct { char c[size]; } y; where size is a run-time value
for hopefully obvious reasons
As for the default libc impl, since it expects a few uncommon c functions, or for example vsnprintf, which is critical for correct aml, it seemed strange to exclude some just because “but memcpy is so common everyone has it lol”
ah so it's a gcc extension
Not obvious, I would have used it for my IPC memes
its an optional part of the standard, though
you shouldn''t
alright, I'm considering switching to clang now
VLAs are bad
But instead I just call alloca

Without the default libc impl it was insanely annoying to port because they had to be propagated via some header anyway, like uacpi cannot just include <string.h> and other headers blindly because they aren't freestanding
(Though malloc and free in the same call context get optimized away by the compilers anyway)
only sometimes
but yeah
new and delete are nowadays common to optimize in C++ compilers
yeah same thing as malloc/free
in some cases, usage of runtime-allocated memory by containers such as std::vector<int> can even be optimized away
it's awesome
and constexpr is awesome too lol
yeah
I think new gets optimized to malloc, which gets optimized away?
no they are specially known to the compiler on their own

yeah
But like C++ optimizations are insane sometimes
you can make any other function work like that too
and I don't think that they are changed to calls to malloc
because operator new is its own thing etc. etc.
and exceptions, obviously. exceptions
oh yeah that too
the prototype is standard, you can just extern it out
I understand your argument for vsnprintf and stuff, but for memcpy its silly
it just is
some functions should be excluded out of the define everything yourself mantra out of commoness and utility
So I should blindly extern stuff by default, but only for some because they're kinda common, and compilers might call them anyway, but not for others because idk
yes
of course if you put it that way it is easier to make it sound like a coocoo idea
uacpis default memcpy outperformed managarms PhD memcpy on lto O3

anything put in clown terms will sound like a circus
the memcpy family
cant you just make the uacpi memcpy a weak function or something if youre not happy removing it
MSVC, LTO builds, static libraries
so like memcpy, memmove, memset, memcmp
bruh
those 4 are required by gcc and clang to exist
Perhaps
perhaps definitely lol
I don't think I have seen a single memmove call ever tho
its mostly just memset/memcpy
i am just saying
i think i might have seen it in rust generated code? idk though
Why
because everyone has them
they are definitely required by clang and gcc, because they always emit calls to them
and you cannot stop that
given 99.9% of people use gcc/clang they will have them, also you already depend on them implicitly on those compilers
True
but what about msvc 
and even people not using gcc/clang, frankly, will have them
But its not true for all mem stuff, literally no code in uacpi out 30k lines generates implicit memcpy calls
like what OS kernel does not have those basic functions that requiring them is absurd?
They might have their own versions of them
you have no way of guaranteeing this
on gcc/clang
as soon as you do *a = *b where a and b are structs, gcc will generate implicit calls
if the structs are large enough
also, GCC will convert loops to memcpy
Yes but its anecdotal evidence not many people will have those
or if the structs have a variable sized member
Referenced implicitly
anecdotal evidence != does not depend on...
Never said that tho
i am stressing the point
Im just saying that the cases of implicit memcpy references dont magically appear out of nowhere
of course not
They're caused by specific code, e.g. large struct copies etc
but you are also not the compiler
a future version may decide that code that rn does not generate those calls, will
ucc when? 
and it would be legal, because they are already stating they depend on those 4 functions
the full list of functions that LLVM will call is here: https://paste.sr.ht/~pitust/5642cb3fb4fb3627069670cd5ec39197ce085d33
or may call, at least
Lily-CC mode that does memcpy for everything >4 bytes? 
Like I agree in general
I dont wanna hear about uacpi bugs because of shitty libc beginners make
for the record it's no mystery, it's very clearly documented in both gcc and clang manuals
I dont think its such a problem to have a built-in memcpy
the full list isnt that clear i think?
but yea
really? that's a bad reasoning. there are a million other ways you can screw up wiring up uACPI and what you're thinking of is a bad memcpy?
if someone can't do a memcpy they should not touch uACPI with a 10ft pole
jesus christ
I've seen some horrors that people made when uacpi required every libc dependency from the host
yeah like lmfao
the beginners you complain will fuck up stuff about already do with the rest of things
clang will memmove if:
- you use memmove, __builtin_memmove, __builtin_memmove_chk
- you use bcopy or __builtin_bcopy
this is just a bad reason
But your reasons are: its common, and your build system cant handle headers well
Those aren't that great either
True I guess
in my opinion, having uACPI define its own libc functions by default is good because
- there might be bugs in the kernel's definitions.
- the kernel might call them something else, like
RtlCopyMemory, or just not at all have them.
a counterpoint to 1) is that it's a kernel skill issue and the kernel should fix it. or that the uACPI libc functions might compile tojmp memcpyetc. anyways (except this doesn't happen if-ffreestandingafaik)
it's common, it's what most normal sane in the head people want, every has those functions, yes, dealing with defines and headers is annoying in other non-make/cmake/etc build systems
and you can still have the old behaviour as opt-in rather than opt-out
- is not a thing for the record, memcpy is required by all language runtimes existing out there, and modern C compilers straight up require them to exist
any good build system would make it easy to add an include path?
I've found that only the most naive implementations of stuff like memset and memcpy will compile into jumps on GCC
imho "API surface minimalism" is not a good goal