#uACPI - a portable and easy-to-integrate ACPI implementation
1 messages · Page 50 of 1
and i686 seems useless as well
i might support compat mode for fun tho
for userspace
also maybe i should make a thread as to not pollute uacpi
which will also make me more embarrassed to abandon it
great idea
it is kinda pointless to support non-64-bit x86 as a platform, since realistically anything nowadays supports 64-bit x86, even the shitty notebook CPUs
compatibility mode is kinda cool though, i think i've see only one (hobby) kernel support it
I think pmos supports it
yes yes exactly
Dunno of other
i couldn't recall the name lol
Thanks for your support
yeah compat mode seems neat
yeah probs
but x86 is a legacy arch, don't you know, and next year everyone will run riscv
on the other hand its nice as you can run old stuff but then there is also the non-nice thing about it which is that it allows new apps to be published as 32-bit binaries (which sucks because on eg. linux you need a ton of redundant 32-bit libraries along with 64-bit ones for eg. steam)
I'd much prefer if it was not a thing and instead if you want to run old apps you would use an emulator
sure bro

My take on this is that 32 bit archs with <= 1gib of memory make some sense
Like, even if you have less than 1gb memory having native 64bit regs is still useful
i think the point is more that 32-bit targets with>1GB of ram make little sense to target
Because only a few of those targets will exist... When you get to that amount of RAM it makes no sense to limit yourself to 32-bit for any reason other than backwards compatibility... Which isn't exactly something x86 struggles with.
As much as I'd love the idea, RISC-V isn't as mature and has not reached the levels of performance and efficiency necessary yet.
i meant like neat technically. i agree with you on the general issues lol
FWIW this is why BadgerOS doesn't bother supporting 32-bit at all: If you're going to have an MMU in 2025, you might as well make it 64-bit.
i dont plan to support compat mode or 32-bit stuff on imaginarium either
on devices with low amount of ram, 64-bit is a fair bit of memory wasted from pointers
Yes
But consider also: MMUs are physically huge
Because they consist basically mostly out of a kind of cache called the TLB
Any chip that can afford an MMU by this definition will already be aiming for higher amounts of memory.
At least nowadays
I mean you would size the TLB roughly.according to the amount of ram you have
You could
But you need to keep in mind that new chips being made get an MMU if they're made for running a modern OS, and don't get one if they're for microcontrollers
there's a middle ground between microcontrollers and full desktop/server capable hardware
Nobody (relevant) is making chips with an MMU that don't have at least a GiB of RAM because an MMU only adds value if you want to run a paging operating system on it, which, in turn, is only useful if you can't properly or cleanly achieve your goal without such an operating system.
And while Linux lets you run with shockingly small amounts of RAM, relatively few such configurations are actually used that do have an MMU.
well that's just not true 😉
i can immediately name one this is untrue of
retrobsd (2.11bsd which has paged vm) was ported to several such microcontrollers, one of which something in the microchip pic32 series
smart watches for example typically have 500mb-2gb ram or so and invariably run full blown operating systems, sometimes in 32-bit
Yeah but microcontrollers with 512KiB of RAM are never gonna have page-based virtual memory
Wait WTF no I'm wrong
But how does a PIC32 with 512KiB of RAM need an MMU?
Granted it's not a big TLB
was that in reply.to me
Both
And then I found a PIC32 with page-based virtual memory
Though it doesn't seem to be entirely a normal TLB from how I read the datasheet?
I mean my example is also hardware that has full blown virtual memory etc
they're mips, it'll be software refileld
no catch, it's a legitimate and peaceful approach
in what ballpark?
my point was that it's a device where you have virtual memory and low ram, so 32-bit ptrs are desireable
motorola coldfire also doing this
Well it is a catch for a TLB with merely 16 entries. You're not getting away with a lot of fragmented virtual memory regions like Linux would make eventually if your TLB has 16 entries.
Not that Linux runs on 512KiB in the first place
Ngl I keep forgetting about the existance of smart watches and all those stupid smart home devices because I just hate them
yeah fair
by comparison the 486 had 32 tlb entries
honestly
it's certainly a squeeze
Why does my lamp need to connect to a network cmon
Yes but that has hardware refill
I guess hobby os for smartwatch confirmed? 
i mean with mips parts of the va space don't go through the tlb
kseg0/1 are just directly mapped to phys mem
free hhdm
I like how this STM32 datasheet says "fuck it" and marks the entire low 2GiB as reserved in the virtual memory map
Every day I learn more how cursed computers are
Anyway I'm off to speedrun Celeste now
also TIL pic32 is mips, i thought it was a custom arch
I thought it was ARM for some reason
i wouldn't have known but for retrobsd making the news a while ago
but it's not hard to support?
still pointless
and there is a factor of it being cool and getting to run your os on old hardware
and being a sanity check for 64 bit code
32 bit kernels should be banned by the eu
i'm 22
i'm 23
i had my first pc at 9 yo
i think me too
and that was an amd bulldozer
i think the earliest machine i had was a c2d one
if it's the only arch
why does nobody support it?
like it's just a 32 bit code segment in GDT, and a couple of if statements to limit virtual memory to 4GB and to get syscall arguments differently
probably because it's difficult to propagate the 4gb limit everywhere
how so?
also probably because not many people want it
compat mode is pretty useless unless you want to run wine
it was convenient when I was porting my kernel to 32 bit mode
but I guess nobody does it in that order?
tbh i think most people just dont support 32 bit kernel mode
moreso than not doing it in that order
Me when my kernel is only 32 bit 
i did say most not all lol
mood
qemu log and printk log or it didn't happen
[00:00:00.05977] [ info ] uacpi: FACS 0x000000007FBDD000 00000040
EXCEPTION: PAGE FAULT
Accessed Address: 0x7FB7B000
Error Code: PageFaultErrorCode(0x0)
InterruptStackFrame {
instruction_pointer: VirtAddr(
0xffffffff800d858c,
),
code_segment: SegmentSelector {
index: 1,
rpl: Ring0,
},
cpu_flags: RFlags(
RESUME_FLAG | INTERRUPT_FLAG | PARITY_FLAG | 0x2,
),
stack_pointer: VirtAddr(
0xffff80007e8c9790,
),
stack_segment: SegmentSelector {
index: 2,
rpl: Ring0,
},
}
.=- -*+-
**+=-: :+++=.
:++**+ .=++= :+***.
:*+++*++:+***= .:::++*+-===-. -++- -++++:
++***+++**= .=**++#%#%%#*+*%%#*-. +***- :=****:
.+*+++****: :*%%*==++#%%%*++++*#%%#+. -**+-*+++***
:*++*++**: .*%%*=+***##########*+#%%%%= -*++***++**-
-****+++ :#%#%##%%*=--::::-=+#%%%%%#%%* -++******+:
-+***= .#%#%%%%#-::::::::::::=#%%%%%#%* -****+:
... *%#%%#%*:::::::::::::::-#%%%%%#%- ..
#%%%%%*::::.:::::::.::::-%%%%%#%+
#%%%%#-:::::........:::::+%%%%#%+
*%#%%=::::..:=++*+-:.:::::*%#%#%-
:%###-::::-*#%%%%%%#+:::::=%%#%#
=##+:::-*%%%######%%#=::::####.
:#*--=#%%#%%%%%%%%#%%+-:-##+.
=%%#####%%%%%%%%%%##%###%%
.::+#%%%%%%%%%%%%%%#*=--:
.:=+*##%%%##*+=:
~~~~~~~~~~~~~~~~~~~~~~~~~ Kernel Panic ~~~~~~~~~~~~~~~~~~~~~~~~~
~
~ ERROR: panicked at src/arch/interrupts.rs:96:5 with message: page faulted
~
~ frame 0: rip = 0xffffffff800fba5f
~ frame 1: rip = 0xffffffff800118ca
~ frame 2: rip = 0x7fb7b000
~ frame 3: rip = 0xffffffff800d85d1
~ frame 4: rip = 0xffffffff800d8949
~ frame 5: rip = 0xffffffff800d9561
~ frame 6: rip = 0xffffffff800d93fd
~ frame 7: rip = 0xffffffff800d9677
~ frame 8: rip = 0xffffffff800d97a1
~ frame 9: rip = 0xffffffff800dd031
~ frame 10: rip = 0xffffffff8004161a
~ frame 11: rip = 0xffffffff80017d14
~ frame 12: rip = 0x0
~ stopping code execution and dumping registers
~
~ r15: 0x0000000000000000 - rsi: 0x000000000000001B
~ r14: 0x0000000000000000 - rdx: 0xFFFFFFFF80116648
~ r13: 0x0000000000000000 - rcx: 0xFFFF80007E8C8C58
~ r12: 0x0000000000000000 - rbx: 0x0000000000000000
~ r11: 0xFFFFFFFF800045CA - rax: 0xFFFF80007E8C8F78
~ r10: 0xFFFFFFFF80116648 - rip: 0xFFFFFFFF800045CA
~ r9: 0x000000000000001B - cs: 0x0000000000000008
~ r8: 0xFFFFFFFF80116648 - rflags: 0x0000000000000096
~ rbp: 0xFFFF80007E8C8C58 - rsp: 0xFFFF80007E8C8B28
~ rdi: 0xFFFF80007E8C8F78 - ss: 0x0000000000000010
~ cr2: 0x000000007FB7B000 - cr3: 0x000000007FF3E000
~
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
Damn your kernel panic is very cool looking
thx, i wish it solved the problem though
info mem in qemu
so no its not mapped
but then uacpi is not mapping it
do u think its a qemu bug
do u think its a uacpi bug
no
i would first add a tracing function that prints everything uacpi tries to map and what you return to it
this is without if i dont do anything in uacpi_kernel_map
so it is mapping things
just not everything
yes, see my message above
after (temporarily) changing the pagesize to 2mb instead of 4kb now the code hangs at uacpi_namespace_initialize()
and i get these errors because i left pci as unimplemented
u can set the log level to debug if u wanna see where it hangs
how? in uacpi_initialize()?
uacpi_context_set_log_level
i dont think uacpi-rs bindings have that lol
you're probably using a very old version of uacpi as well
are u using the rewrite branch?
maybe it still using the old version where uacpi_initialize takes in a log level
then it should have that function
at least in C
maybe its not exposed to rust
only @rustic compass knows
or just use addr2line or whatever
in which file is that function? maybe its not included in the bindgen sources
context.h
alternatively u can change the default log level in sources
i think config.h has it
yeah this is available - uacpi_set_resources
what
nvm
does debug log every single instruction?
every aml op
depends on how fast your logger is
the last thing it logged
[00:00:44.87663] [ dbug ] uacpi: trying to acquire the global lock from firmware... (attempt 1)
[00:00:44.87853] [ dbug ] uacpi: global lock is owned by firmware, waiting for a release notification...
spinlock?
ah ok its a bug in your kernel api
either spinlock api, or mutex api, or get_thread_id
all of those are set to 1 haha
set to 1?
i dont see a get_thread_id in there
its not in the file u posted
bruh i didnt copy the entire thing
yeah nvm
its because u return 1 for all mutexes
handles collide
anyway just implement it properly
or at least return an incrementing number each time
lol
ok that will probably fix the hang but am i supposed to always use 2mb page sizes in the mapping?
4kb pagefaults
no, your page mapper is just bugged
like i said trace everything uacpi asks to map and the address you return, then compare with the resulting info mem
fixed the mapper
What was the issue?
nothing anymore, thanks
Lol np
dumass forgot a loop and was mapping a single page
Lmfao
how good is 2408331 ops/s
so im getting better results than people that have a better cpu than me
which CPU do you have?
i5 12400f
it's not that slow...
never said it was,
Its pretty decent
2681214/s now
weirdly enough setting the opt-level to anything other than z or s on release profile makes the os bootloop
Something like 12cm. 
works for me on opt level 3

FPGA MFs when they forget to add outputs to their design
what does it mean if im getting UACPI_STATUS_INIT_LEVEL_MISMATCH from this?
nvm i figured it out
i was initing hpet before uacpi
can't wait for the ultra kernel progress thread
when theres at least something to show lol
do it now
Many threads, including BadgerOS', start as basically nothing

i will make one soon
I would recommend setting an explicit goal, as in "progress report thread when XYZ thing is accomplished" (e.g. userspace)
When I make a good log ring 
me when I always rewrite before reaching the goal I set to myself for creating a thread 
the double fault is caused by some "normal" fault handler (e.g. GP or PF) not being installed.
aka give QEMU log
What's at that line?
or not being installed correctly, or something being wrong with the stack
its not a GP or PF, both are installed correctly and work
Then its probably a push to a bogus rsp or something
gib qemu log !!
Holy sgit alnost 50k messages, this project is big as heck
qemu log or it didn't happen
yes but also no, this is just lounge-2
By that measure nyaux is a bigger project
enable qemu interrupt logging and check exactly what exception is being thrown before the double fault
maybe your stack is too small and it can't push the exception?
that doesnt work with kvm
and without kvm the double fault doesnt happen
Use bochs
try to give all the exceptions a specialized stack
hmm ive increased the stack size to 8kb and no more double fault
I would put a stack guard, so you can ensure you never run out of stack accidently
im gonna set it to 16kb to be safe
I think linux pushed it to 32k as more rust code enters the kernel
If your entire kernel is rust maybe consider 64 or even 128
unrelated to the current conversation but my issue with fixed events causing a page fault was due to my shitty implementation of std::function
now pf here, what could this be caused by?
Lol
Bad mapper or any other corruption
also I don't think this is how time works
THIS IS SO WEIRD
Seems to increase monotonically, ship it
man, it sounds like there is something really wrong with your code lol
64kb works
for your own good, put a guard page for your stack and check if that is actually related
maybe your allocator is cooked
maybe its your page allocator
that part might actually be cooked
and again, stack guards are a life saver when it comes to these stuff
it adds extra 5 seconds every 5 seconds
is there any sort of alignment guarantee on uacpi_resource?
anyway imaginarium and zuacpi now can do all resource types that showed up in qemu. will finish and polish this bit up and then push in a few days
Yeah
hows your os going btw
:(
wdym
thats why i started working on stuff in managarm
depressing
i am incapable of creating a project and working on it for extended periods of time
you are a great help to me tho
yep
same
at this point, any second set of eyes is a win in my book
so i don't produce utter garbage code
it’s easier when working on someone else’s code, you don’t have to think about the entire codebase usually
merge yours with mine :)
doesn’t work, mine isn’t written in rust
and i don’t really have a vision for mine but i don’t want to introduce a bunch on shit code lol
i don’t really know what i’m doing is the problem
pretends that i do
i have 0 clue past the surface level
i'm just making shit up as i encounter it
same, but so far it has worked for you, unlike me :^)
luck

Impostor syndrome
you just need to be willing to read linux sources and other crap for hours
at least thats what i do
i despise doing that
how do you make an abstraction layer this annoying to work with
im kinda used to it, at this point its relatively easy to navigate
yes i kinda hate it as well lol
mine isn't either 
yeah but now i don't have anything that would be worth contributing to your kernel
- you use gitlab, ew
natural alignment I assume?
it makes sure its aligned to reg width basically
yeah I was just trying to figure out what the alignment was lol
cause otherwise im having to aligncast
uacpi tries to be sane where possible
my zig code was using align(1) aka c packed just to be as general as possible until I figured it out lol
lol
also zig packed is completely different and I keep having to remind myself that c packed isn't the same when talking with c people lol
yeah GCC packed means remove padding AND assume misaligned
and zig packed means bit packed
yup
anyway i can remove a bunch of code in one spot and change a 1 to a @alignOf(usize) in another now and have zuacpi still work at least which is nice
love removing code
can you guys test my kernel on a high end cpu? i wanna see how much points it can get, most i could get was 3.3m on my i5 12400f
RUST_PROFILE="smol" make run (this has the most optimization)
without ovmfs its slower
doesnt compile for me
why?
Updating crates.io index
error: failed to get `flanterm-sys` as a dependency of package `chronos v0.1.0 (/home/cat/prob/chronos/kernel)`
Caused by:
failed to load source for dependency `flanterm-sys`
Caused by:
Unable to update /home/cat/prob/chronos/bindings/flanterm-sys
Caused by:
failed to read `/home/cat/prob/chronos/bindings/flanterm-sys/Cargo.toml`
Caused by:
No such file or directory (os error 2)
make[1]: *** [GNUmakefile:30: all] Error 101
make[1]: Leaving directory '/home/cat/prob/chronos/kernel'
make: *** [GNUmakefile:70: kernel] Error 2
did you clone it recursively?
oh
also you will get pci unimplemented errors which cover the screen so you will have to check the terminal debugcon prints for the points
i implemented pci now but haven't pushed it yet
rust-lld: error: too many errors emitted, stopping now (use --error-limit=0 to see all errors)
warning: `chronos` (bin "chronos") generated 1 warning
error: could not compile `chronos` (bin "chronos") due to 1 previous error; 1 warning emitted
can you show rust-lld errors?
hm thats wewird
are you using nightly rust?
that shouldnt matter for linking but
jic
i just ran make run, and it asked me to add nightly so i did
what distro are you using?
fedora
im not sure why this is happening...
@torpid root can u test it?
i like how u start pinging random people to test it for you while i offered to test it for you, you still haven't provided the image and a command line
yeah that is much easier
he's not random, I know him and he runs fedora
i'll provide the iso and commandline in a sec
well, yeah, random as in out of nowhere
that's what i seemed like from my point of view
qemu-system-x86_64 -M q35 -cpu host -debugcon stdio -boot order=d,menu=on,splash-time=0 -cdrom chronos-x86_64.iso -enable-kvm -m 2G
^
you don't compile your dependencies with the same code model as your kernel
add -mcmodel=kernel
i fixed it btw, you can clone again
hel syscall compat kernel 
No that's not a good idea lol
[iretq@raptor ~]$ qemu-system-x86_64 -M q35 -cpu host -debugcon stdio -boot order=d,menu=on,splash-time=0 -cdrom chronos-x86_64.iso -enable-kvm -m 2G | grep -a '\(avg'
[00:00:00.00428] [ info ] uacpi: successfully loaded 1 AML blob, 1705 ops in 0ms (avg 3837030/s)
[00:00:00.00364] [ info ] uacpi: successfully loaded 1 AML blob, 1705 ops in 0ms (avg 3837160/s)
[00:00:00.00299] [ info ] uacpi: successfully loaded 1 AML blob, 1705 ops in 0ms (avg 3864224/s)
[00:00:00.00323] [ info ] uacpi: successfully loaded 1 AML blob, 1705 ops in 0ms (avg 3421413/s)
[00:00:00.00305] [ info ] uacpi: successfully loaded 1 AML blob, 1705 ops in 0ms (avg 3519520/s)
[00:00:00.00307] [ info ] uacpi: successfully loaded 1 AML blob, 1705 ops in 0ms (avg 3814513/s)
[00:00:00.00338] [ info ] uacpi: successfully loaded 1 AML blob, 1705 ops in 0ms (avg 3823307/s)```
13700K
nice
thanks
didnt you have a python auto-launcher
idk lmao
maybe
i do vividly remember making something like that
yeahhh right, we were debating on how to make it get the most cpu time possible
yeah
i define this function in rust the same way as the kernel api functions
what am i doing wrong?
how is C supposed to know what chronos_memcpy is
how does it know with kernel api functions?
it includes the header where they are declared
also just call this function memcpy and it will be picked up by uacpi automatically
because it uses __builtin_memcpy
and you also get memcpy from rust compiler builtins so you likely don't need to make one
use bindgen;
use cc;
fn main() {
let mut b = cc::Build::new();
b.files([
"uacpi/source/default_handlers.c",
"uacpi/source/event.c",
"uacpi/source/interpreter.c",
"uacpi/source/io.c",
"uacpi/source/mutex.c",
"uacpi/source/namespace.c",
"uacpi/source/notify.c",
"uacpi/source/opcodes.c",
"uacpi/source/opregion.c",
"uacpi/source/osi.c",
"uacpi/source/registers.c",
"uacpi/source/resources.c",
"uacpi/source/shareable.c",
"uacpi/source/stdlib.c",
"uacpi/source/tables.c",
"uacpi/source/types.c",
"uacpi/source/uacpi.c",
"uacpi/source/utilities.c",
])
.includes(["src/"])
.includes(["uacpi/include/"])
.define("UACPI_SIZED_FREES", None)
.define("UACPI_OVERRIDE_LIBC", None)
.pic(true)
.flag("-ffreestanding")
.flag("-nostdlib");
if std::env::var("CARGO_CFG_TARGET_ARCH").unwrap() == "x86_64" {
b.flag("-mgeneral-regs-only");
b.flag("-mno-red-zone");
}
b.compile("uacpi");
let bindings = bindgen::builder()
.use_core()
.wrap_unsafe_ops(true)
.derive_default(true)
.derive_debug(true)
.header("src/wrapper.h")
.clang_arg("-Iuacpi/include/")
.generate()
.expect("Unable to generate bindings!");
let out_path = std::path::PathBuf::from(std::env::var("OUT_DIR").unwrap());
bindings
.write_to_file(out_path.join("bindings.rs"))
.expect("Unable to write bindings!");
}
this is my bindgen
do you really need sized frees with rust?
hm ok
and what do you do with the UACPI_OVERRIDE_LIBC
bruh
and where is that file?
make it yourself, but infy said it's redundant
#ifndef __smp_store_release
#define __smp_store_release(p, v) \
do { \
compiletime_assert_atomic_type(*p); \
__smp_mb(); \
WRITE_ONCE(*p, v); \
} while (0)
#endif
#ifndef __smp_load_acquire
#define __smp_load_acquire(p) \
({ \
__unqual_scalar_typeof(*p) ___p1 = READ_ONCE(*p); \
compiletime_assert_atomic_type(*p); \
__smp_mb(); \
(typeof(*p))___p1; \
})
#endif
@flat badge any idea why linux uses a full barrier for both of these instead of smp_rb() for one and smp_wb() for another?
i looked at the commit that introduced this but no one asked this question
so what should the default name be? just memcpy or uacpi_memcpy
oh
i might steal this

mine is better
i guess this is a generic impl that is overriden by archs?
yes
yeah on x86 it compiles to a compiler barrier
ah the question was not why is it not a acquire/release barrier but why is it not smp_rmb / smp_wmb
yea
i dont think they even have that
that's why they use the strong full memory barrier ;D
so in linux terms smp_rb means no reads can be reordered with this read?
but not writes
but compare for example on RISC-V where linux lowers smp_rmb() to RISCV_FENCE(r, r)
but acquire to RISCV_FENCE(r, rw)
so the C model is stricter, it doesnt even have a way to express an smp_rb or smp_wb
a read-read barrier is not strong enough because that'd make it possible to reorder
mutex.acquire()
x = 42
mutex.release()
to
x = 42
mutex.acquire()
mutex.release()
yeah makes sense, so with smp_rb, stores can be reordered in any way
so with stdatomic you could really only express smp_rb_readwrite(), not smp_rb{_read} like linux
the C model is different not finer/coarser
seems coarser to me
a read-rw barrier is not equivalent to acquire either
you just cant express this
read-rw is stronger than acquire
why?
because acquire only synchronizes with release barriers on other CPUs
and not with random relaxed accesses
well if we talk about a std::atomic_thread_fence
with ACQUIRE
it would mean the same
no
how
i dont understand the difference
the Linux model is phrased in terms of what accesses before a barrier can be reordered with what accesses after a barrier
right
the C11 model is phrased in terms of when one access on one CPU happens before another access on a different CPU
same thing just a different angle?
happens-before/sequenced-before is just lawyer speak for the exact same thing if i understand this correctly
doesnt this basically say they can't be reordered just using a more abstract/formal language?
btw im also watching a talk by Paolo Bonzini and he had this
okay, i think you're right that in the case of fences, r-rw is equivalent to acquire
and rw-w is equivalent to release
I read this probably 100 times
and still dont understand what this is trying to say
but there is no way to emulate r-rw with smp_rmb() and smp_wmb()
yeah, u must use smp_mb
why rw-w and not w-rw?
otherwise you can again come up with a silly mutex example
otherwise
mutex.acquire()
*x
mutex.release()
can be reordered to
mutex.acquire()
mutex.release()
*x
wait what do u encode in the order of these letters exactly
rw-w = no read/write before the barrier can be reordered with a write after the barrier
okay and r-rw means no read/write after the barrier can be reordered with this read?
r-rw = no read before the barrier can be reordered with a read/write after the barrier
because this sounds kind of inverse from what you're saying
that's exactly what i'm saying?
except that i'm talking about linux style barriers while that paragraph talks about acq/rel load/store
nvm ill have to re-read this probably 50 times to parse it completely
did u understand this part?
bonzini based
yes, what is your question here?
i knew him from the smalltalk community long before i encountered him elsewhere
he is, still one the most active qemu maintainers
i think its trying to say that code before release cannot be reoredered with code after acquire or whatever
it means that if
thread 1:
x.store(42)
thread 2:
x.load()
and thread 2 loads the 42 written by thread 1, then everything that was ordered before the store in thread 1 is also visible to thread 2 after the load
yeah makes sense when you put it like that
https://www.youtube.com/watch?v=NE73iPMpzj4 its this by the way
C gained its popularity because of its flexibility and the ability to function as a "high-level assembler". However, due to the growing complexity of programming both languages and code bases, compiler writers need to extract more and more information from the source code, causing behavior that the programmer did not expect---the dreaded "undefi...
pretty good talk
when does uacpi_kernel_uninstall_interrupt_handler get called?
uacpi_state_reset
@flat badge in this example, why is acquire_release on the load()'s not enough? Wouldnt it guarantee that stores before cant be reordered after this one
or is the idea that there isnt a synchronization point between them
like this makes no sense, because surely if x was a normal variable then it would be guaranteed to be false after an acq_rel load, no?
like its this sentence i dont understand
if this was the case, then
// t1
foo.store(123, relaxed);
bar.store(321, release);
// t2
bar.load(acquire);
assert(foo.load(relaxed) == 123); // can fail beacause not seq_cst??
why wouldn't bar's release be enough here
actually if this is true then they arent equivalent
because the c++ page simply says "no reads or writes in the current thread can be reordered before this load", it doesnt say anything about reordering with a write specifically
or e.g. why a thread_fence(ACQ_REL) wouldnt be enough between each load and store, they could even be relaxed
iirc the happens-before relation that acq rel creates only applies to non atomic vars
foo is atomic, so its store doesn't happen before bar's
hm so my example can literally fail?
yeah
lol
if this is the answer then i understand why u need seq_cst, but i hope its not true
it also means that mutex acquires need to be acq rel instead of just acq since otherwise lock(a); lock(b); can be observed as lock(b); lock(a); on other threads
this can't fail
I guess monkuous is wrong then
but in this snippet there is no such synchronization
thread 2 doesn't load acquire any value written by thread 1 (or vice versa)
that's why you need seq_cst
yeah but what if i add an acq_rel barrier in the middle for example
x.store()
fence(acq_rel) <--- does nothing here
y.load()
acquire prevents reordering earlier loads
the store cannot be reordered after the fence
there is no earlier load
the load cannot be reordered before
release prevents reordering future stores
and there is also no future store
so acq_rel does nothing here
well the docs dont say that
no reads or writes in the current thread can be reordered before this load
and no reads or writes in the current thread can be reordered after this store
it mentions both
by "this load" this summary means "an earlier load" (or the in the case of a load-acquire instead of a fence, the load done by the load-acquire)
A load operation with this memory order performs the acquire operation on the affected memory location: no reads or writes in the current thread can be reordered before this load.
this is the entire text
pretty sure this load means the load itself
there is no load acquire here
for acq_rel it says No memory reads or writes in the current thread can be reordered before the load, nor after the store.
acq_rel is equivalent to acquire + release
but this is still different from what the cppref page says, no?
yours seems lighter as it allows reads to be reordered with reads
if you make it
T1:
x.store(false, release)
a = y.load(acquire)
T2:
y.store(false, release)
b = x.load(acquire)
you can still get a = true, b = true
there is simply no synchronization at all here
i agree with this one, since the store is before the acquire
acquire means nothing below can be reordered above
and
T1:
x.store(false)
fence(acq_rel)
a = y.load()
T2:
y.store(false)
fence(acq_rel)
b = x.load()
is equivalent to the above
are these relaxed?
so what you're saying is that the fence is a no-op for some reason
but why
would it also be a no-op if load and store were in a different order?
ok i see what u were going for
well, if you change the order in one of the two threads
what about this tho?
i think your definition is lighter still
what you're saying is
foo.load(acquire);
print(x); // may be reordered above the load
if i understood you correctly
but cppref seems to disagree
linux docs agree with cppref
no
isn't it the other way around? acq prevents reordering future loads to before the acq, rel prevents reordering past stores to after the rel
it is, but it's both loads and stores
i think what you're describing is the equivalent of smp_{rb,wb} which are not acquire/release
I'm pretty sure the cppref summary is incorrect here and the formal model doesn't care about stores wrt acq and loads wrt rel
what makes you think that I stated that?
that is not the situation of the snippet that you posted above
where the load is after the store, not before the store
yes it does
otherwise stuff could be moved out of a mutex
r-rw = no read before the barrier can be reordered with a read/write after the barrier
wait wrong one
rw-w = no read/write before the barrier can be reordered with a write after the barrier
this implies a read inside the barrier can be reordered with anything
r-rw (i.e. acquire) prevents #1217009725711847465 message
then it implies that stuff inside the barrier can be reordered with a read after
the load before the barrier (foo.load) cannot be reordered with anything (the print) after the barrier
which doesnt make sense
are you aware that if you want to replace load-acquire with a fence, the fence has to be after the load?
what you said is "reads/writes before the barrier can be reoredered with reads after the barrier"
and likewise if you replace store-release with a fence, the fence has to be before the store?
well yeah
that i understand
if you write this #1217009725711847465 message in terms of fences, it is
foo.load(relaxed);
fence(acquire)
print(x);
and the r-rw fence prevents reordering
yeah, so it cant be reordered
ok so r-rw = no read before the barrier can be reordered with a read/write after the barrier AKA for the acquire barrier, you said that writes before the barrier can be reordered with anything after the barrier
correct?
yes
nvm
there is no write before the barrier here
the semantics are literally what i wrote though lol
No, i meant literally what i wrote
and not anything else
then your semantics are actually stricter
because cppref doesnt mention anything about the operations before the barrier
it only mentions that operations after cant be reordered
that doesn't make any sense
ordering is always before barrier vs. after barrier
you can't say something about the ordering on one side only
what
and cppref also says something about both sides
it says stuff after cannot be ordered before
what does "before" mean?
it doesnt mention what can happen to operations before the barrier
because its not relevant for the acquire barrier i guess
there is no such thing as a "nothing-r" or "nothing-rw" or "nothing-w" barrier
because it would be a no-op
well thats not true
operations before the acquire barrier can happen after the barrier
the only limitation is operations after barrier must appear to happen after the barrier
stuff before the barrier is not synchronized in any way tho, why is it important to define it
(1) that's not how it works in the C11 model and (2) that wouldn't make any sense
thats literally how its spelled out A load operation with this memory order performs the acquire operation on the affected memory location: no reads or writes in the current thread can be reordered before this load.
mentions absolutely nothing about before the barrier operations
can be reordered before this load
that's "before"
if you leave out this sentence, the CPU can delay the barrier indefinitely
yes, but its talking about the operations after the barrier
ok, why?
actually that's not true, it could take the barrier immediately upon program start
it could move it into the past, not into the future
then everything after the barrier stays after the barrier and there are no memory accesses at all before the barrier
they're there, we just dont talk about them
the "until the heat death of the universe" scenario applies to barriers with empty "after" side
Barriers where one side is empty are meaningless in any case
fences aren't loads or stores, they effectively raise the ordering of other loads or stores (depending on the order of the fence)
by your definition, a program that starts with an acquire barrier instruction isnt allowed to exist
because there isnt anything before it
bullshit
yeah i get that
what im saying is for the purposes of the acquire barrier, the only loads or stores that matter to us, are those that happen after the barrier
i've yet to understand why thats not the case
any read before the acquire fence (this set is empty if the acquire is at program start) cannot be reordered with any read or write after the acquire fence (the entire program)
why read specifically?
because that's the freaking definition
where
an acquire fence raises the ordering of all prior atomic loads to acquire. that's all it does
you cited it 10x already from the cppref summary
no reads or writes in the current thread can be reordered before this load
"this load" mentions no prior loads?
this load refers to
- the load done by the load-acquire
- in case of fences: to all previous loads
so then "any read" in your definition means "this read" for a single atomic (non-fence) load?
the cppref summary is imprecise because it doesn't distinguish fences and load-acquire
acquire fences are stronger than load-acquire
yeah i think this is where we misunderstood each other
because acquire fences apply to all previous loads while load-acquire applies to exactly one load, the one done by the load-acquire
likewise, release fences apply to all future stores while store-release applies to exactly one store, the one done by the store-release
so then
any read before the acquire fence (this set is empty if the acquire is at program start) cannot be reordered with any read or write after the acquire fence (the entire program)
for one load-acquire can be transformed into
an acquire load cannot be reordered with any read or write after the acquire load
correct?
yes
lmfao
yeah sorry i misunderstood what u were trying to say
this is what i meant this whole time as well
because fence vs load confusion
btw by "any read before the acquire fence" do u mean only atomic reads, or do non-atomic ones also get promoted to acquire
so basically non-atomic reads before an acquire fence can be reordered after the fence?
if they're after the last atomic load the fence applied to, sure
ah right, if they're before that still works fine
or well acquire doesn't prevent reordering things before it to past it so any read can be reordered to after an acquire fence
but the same would be true if you replaced the fence with an acquire load
but yeah C's fences are only useful for lifting relaxed loads/stores into acquire/release ones
so they're basically a bulk convert previous atomic op ordering into something else
or well, following for release i guess
yep
Tbf that's how fences are usually implemented in CPUs
x86 is a bit different, because the default semantics are already read acquire store release (or well, technically stronger but still)
And cmpxchng is basically seq_cst
Or well, the locked version
hm
Mfence is used for seq_cst operations
And I think sfence and lfence are only useful for the non-temporal moves
Since those bypass tso
i guess linuxes ones actually guarantee that this applies to all previous operations, not just the atomic ones
(3) Read (or load) memory barriers.
A read barrier is an address-dependency barrier plus a guarantee that all
the LOAD operations specified before the barrier will appear to happen
before all the LOAD operations specified after the barrier with respect to
the other components of the system.
A read barrier is a partial ordering on loads only; it is not required to
have any effect on stores.
Read memory barriers imply address-dependency barriers, and so can
substitute for them.
[!] Note that read barriers should normally be paired with write barriers;
see the "SMP barrier pairing" subsection.
If you really wanna have fun I think aarch64 has a format memory model definition, including litmus tests and all the fun stuff
How do you test that stuff
updated and is about ~100k-200k faster on my pc
may i request to be put on the leaderboards?
There are ways to formalize the memory model semantics and then write a test to check if that condition can happen under that memory model
Lemme find the aarch64 one
Yes, just give me a link, short description, and the CPU
Thanks
I am not gonna pretend like I ever looked at it too deeply, but I just think it's pretty cool
https://github.com/BUGO07/chronos
another x86_64 kernel held together by duct tape, made in rust
~3M points on i5 12400F
~3.7-3.8M points on i7 13700K (should be ~4M now)
Can you give the full number
What's cool about this is that once they defined all the semantics they want, they can test them against the models of arches that also have their semantics formalized, so they can ensure that they use the correct instructions to implement their barriers to keep the Linux semantics correct
Yeah
Thx, ill take a look
3864224 on i7 13700K
im the only one with 3.xxM on the leaderboards
i see you're using talc, how do you do page allocations with that?
3892685/s
basically the same
It's about 2kk on 5900X
check memory/vmm.rs
I sleeb
Or don't, I get an overflow error when booting on my laptop on debug builds
I don't think he does
does what?
page allocations
You're right that this affects only atomic ones, but that's more an artifact of how C works
For non atomic vars, C doesn't even guarantee that they are written by a single CPU operation
So it can't make any synchronization claims about them
I also don't think that linux's memory model does
Linux has READ_ONCE for this purpose
On existing ISAs, properly aligned single instructions accesses are always relaxed atomics
except that a normal C load is not because it is subject to optimizations ofc
do acquire fences even apply to normal reads and writes?
Depends on what you mean by apply
have any impact on them
#define __READ_ONCE(x) (*(const volatile __unqual_scalar_typeof(x) *)&(x))
this definition of read_once is somewhat unsound aiui
yeah the linux memory model is a mess in general
better than nothing for sure
and it makes a lot of assumptions about what GCC does
but its invalid regardless
the VERY fun thing is that while a volatile read occuring is a side effect
the value of it is not a program input (afaik)
essentially you need a atomic var for the synchronization, but the fence still orders the non-atomic accesses around it
otherwise, non-atomic accesses could be moved out of mutexes etc
the fence there only works together with an atomic write
i.e. it attaches to a surrounding atomic
if you have
b.load(relaxed)
fence(acquire)
a.load(non_atomic)
then the acquire fence can potentially make the load of b synchronize with a store release
So in the C memory model, how would you even synchronize something like:
atomic_load(x, acquire);
// unrelated data
memcpy(...);
// reload to make sure the state is the same
atomic_load(x, acquire);
With this memcpy can be reordered to after the load since its not an atomic store
atomic_load(x, acquire); doesnt prevent writes from going below it
only from reads and writes coming above it
Read the question
Yeah that's the only solution I guess
Have an unrelated atomic store
acquire is not strong enough
you need sequentially consistent fences here
Would a seq cst fence be enough to prevent memcpy reorder?
Lol
no nvm i think it would so long as that second atomic load is present maybe?
And do u actually need a fence or is it enough to keep seq cst on both loads
Yeah no idea
the answer is "ask a model checker"
Too skill issued
this doesn't even work if you replace the memcpy by relaxed atomics
btw
the last load-acquire doesn't gurantee anything
Well it guarantees synchronization with a previous store, but doesn't prevent reordering with memcpy that's true
even if you treat the memcpy as atomic accesses, the final load acquire would only guarantee that following loads are more recent than the value loaded by the load-acquire
Linux code used two smp_rb barriers with relaxed loads here btw
what does smp_rb do?
it's a read-read barrier
ah
Prevents read recorders yeah
so dmb ish on arm
and a compiler fence on x86
based on webkit headers: https://github.com/WebKit/WebKit/blob/main/Source/WTF/wtf/Atomics.h
Yeah so u would have to release store for each byte lmfao
C memory model is not less stupid that Linux tbh
at least it maps somewhat better to hw
the linux memory model is stupid beyond imagination
it's also only held up by hopes and dreams that GCC doesn't change its behavior
Lol
they probably won't because of linux
Well they have a .cat file
and then clang comes along
lol
clang already handles volatile access differently than GCC
Btw you suggested a seq cst barrier initially, but aren't these barriers basically used to elevate the ordering of surrounding atomic operations? Like what would be the difference of a pair of seq cst barriers vs just marking both loads seq cst
And iirc seq cst doesn't play nice with non seq cst stores from other places or whatever
There was some weird interaction there
So would all stores/loads to that var have be seq cst
?????
that's stupid
the correct way to do a seqlock is:
doesn't seq-cst mean no reads or writes before can be moved after that operation and no moves after can be moved before that operation?
Well acquire only prevents the following stuff to be reordered with the atomic acquire
That's just how it works
why is atomic stuff so complicated to reason about all of a sudden
real
seq.load(acquire)
if (dirty)
retry
data1.load(relaxed)
data2.load(relaxed)
data3.load(relaxed)
fence(acquire)
seq.load(relaxed)
if (mistatch)
retry
Because modern hw is hard 💀
hw was a mistake
it's not like any of it applies to x86
it's just compiler optimizations here
well you posted a seqlock
It's not, rust docs explain it perfectly
how does this prevent data1..3 loads being reordered to after the last seq.load?
it does not
the acquire fence guarantees that the last seq.load is at least as new as all the data reads
yes it is
lol
ahh i see
thats like the point
why have we spent 5 hours talking about atomics here
if we could just read the rust docs
Ikr
ahh yes nothing like rust
the rust docs have an extremely oversimplified summary of it
i feel like atomics are one of those "if you think you understand it, think again" things
can u elaborate
the entire reason for this convo is that they are very complicated
i was refering to this
i'm sooo glad i don't understand any of it, i simply don't have to worry about it :^)
how would you synchronize the above code with the C model
i think ill never understand atomics
if you want a pure standards-compliant version: use an atomic_memcpy where every byte load is a relaxed atomic load and put an acquire fence between the atomic_memcpy and the last load
Is the reason this works because data load reordering past the check doesnt matter?
it does i think
atomic_memcpy is not a thing tho i think?
my understanding is that this code is invalid
atomic_store(dst[i], src[i], release)?
relaxed
and a release barrier after?
and more like dst[i] = atomic_load(src[i], relaxed)
or alternatively
then atomic_fence(acquire) and another load
the acquire fence will prevent the load after it from being reordered before any of the relaxed loads in the memcpy
hm
__attribute__((naked)) void atomic_memcpy(char* dst, const char* src, size_t len) {
// don't tell clang
asm("jmp memcpy");
}
lmao
but clang would have to be able to see the atomic loads lol
no it doesnt
otherwise it cant prove anything
write side:
s = seq.load(acquire)
seq.store(s + 1, relaxed)
fence(release)
data1.store(relaxed)
data2.store(relaxed)
data2.store(relaxed)
seq.store(s + 2, release)
clang has to assume that inline asm performs an unknown sequence of AM operations
with this it's like atomic_memcpy was implemented in another TU and there's no LTO involved
which can include an atomic memcpy
it doesn't know anything about what it does, so it has to assume the "worst"
which in this case works in our favor
just attribute(alias_of)
actually the first load can be relaxed i guess
or .globl atomic_memcpy; .set atomic_memcpy, memcpy
actually idk if that works across shared objects
dont think so? otherwise there isnt a happens-before with the release
that could be reasoned about
with which release?
asm cannot be
the last seq release by the previous writer
the acquire synchronizes with the last update of the seqlock
but if that's not needed because the entire contents are overwritten anyway, it can be relaxed
depends on what exactly you do with data1 etc
if you rmw it, it needs to be acquire
what does seq have to do with the contents
if you just store independent data, it can be relaxed
then stores can be reordered to before the release
no, stores can not be reordered before release barriers, that's the point of release
An atomic operation A that is a release operation on an atomic object M synchronizes with an
acquire fence B if there exists some atomic operation X on M such that X is sequenced before B
and reads the value written by A or a value written by any side effect in the release sequence headed
by A.
this doesnt seem like it supports this being correct
thats not how that works tho
data1.store() can be reordered, a release above doesnt protect it
what exactly are you concerned about?
reordered to where?
to after the relaxed load
a preceeding release fence prevents reordering, yes
it's not clear to me what you mean, can you post the whole listing?
where can what be reordered to?
s = seq.load(relaxed)
seq.store(s + 1, relaxed)
fence(release)
data1.store(relaxed)
can be reordered to
s = seq.load(relaxed)
data1.store(relaxed)
seq.store(s + 1, relaxed)
fence(release)
the fence turns data1.store into a released store
yeah this can't happen
seq is a memory access before a released store
so seq cannot be reordered to after data1
that the data loads can be moved past the fence
and then the seq load can be moved before them
the only reorderings that are possible are reorders of the data1 vs data2 vs data3 stores
wait I always forget whether these fences apply to previous or next operations
release applies to next, acquire to prev
whats the intuition here?
if that program was in fact correct, then this would not panic: ```rs
fn thread1() {
loop {
let data = sh.data.0.load(Ordering::Relaxed);
fence(Ordering::Acquire);
let seq = sh.seq.0.load(Ordering::Relaxed);
assert!(data <= seq);
}
}
fn thread2() {
let mut n = 0;
loop {
sh.data.0.store(n, Ordering::Relaxed);
fence(Ordering::Release);
sh.seq.0.store(n, Ordering::Relaxed);
n += 1;
}
}
and it does panic
that is not what the program does
hmm
is it not?
am i missing fences?
ah wait yeah i am
okay
a release fence makes all future stores release stores, and thus any previous memory ops cannot be reordered past the first atomic store after the fence
an acquire fence makes all past loads acquire loads, and thus any future memory ops cannot be reordered to before the last atomic load before the fence
std::thread::Builder::new()
.spawn(move || loop {
let seq1 = sh.seq.0.load(Ordering::Acquire);
let data = sh.data.0.load(Ordering::Relaxed);
fence(Ordering::Acquire);
let seq2 = sh.seq.0.load(Ordering::Relaxed);
if seq1 == seq2 {
assert!(data == seq1);
}
})
.unwrap();
let mut s = 0;
loop {
sh.seq.0.store(s, Ordering::Relaxed);
fence(Ordering::Release);
sh.data.0.store(s, Ordering::Relaxed);
sh.seq.0.store(s + 1, Ordering::Release);
s += 2;
}
is this better?
this still panics, btw
loading a number lower than seq1
the write side is still wrong
ah
you want to store s + 1, then s + 2
i am
std::thread::Builder::new()
.spawn(move || loop {
let seq1 = sh.seq.0.load(Ordering::Acquire);
let data = sh.data.0.load(Ordering::Relaxed);
fence(Ordering::Acquire);
let seq2 = sh.seq.0.load(Ordering::Relaxed);
if seq1 == seq2 {
assert!(data == seq1, "loaded {seq1} vs {data} vs {seq2}");
}
})
.unwrap();
loop {
let s = sh.seq.0.load(Ordering::Relaxed);
sh.seq.0.store(s + 1, Ordering::Relaxed);
fence(Ordering::Release);
sh.data.0.store(s, Ordering::Relaxed);
sh.seq.0.store(s + 2, Ordering::Release);
}
``` better?
same result
because the program is equivalent lol
sh.data.0.store(s, Ordering::Relaxed); <--
i just have s cached because there is onlt one writer
oh wait
yeah
nvm
std::thread::Builder::new()
.spawn(move || loop {
let seq1 = sh.seq.0.load(Ordering::Acquire);
let data = sh.data.0.load(Ordering::Relaxed);
fence(Ordering::Acquire);
let seq2 = sh.seq.0.load(Ordering::Relaxed);
if seq1 == seq2 {
assert!(data >= seq1, "loaded {seq1} vs {data} vs {seq2}");
}
})
.unwrap();
loop {
let s = sh.seq.0.load(Ordering::Relaxed);
sh.seq.0.store(s + 1, Ordering::Relaxed);
fence(Ordering::Release);
sh.data.0.store(s, Ordering::Relaxed);
sh.seq.0.store(s + 2, Ordering::Release);
}
ah wait nvm
try to convert my pseudocode to a working program challenge (impossible)
lol
i mean yeah your code is broken so like (/j)
std::thread::Builder::new()
.spawn(move || loop {
let seq1 = sh.seq.0.load(Ordering::Acquire);
let data = sh.data.0.load(Ordering::Relaxed);
fence(Ordering::Acquire);
let seq2 = sh.seq.0.load(Ordering::Relaxed);
if seq1 == seq2 {
assert!(data == seq1, "loaded {seq1} vs {data} vs {seq2}");
}
})
.unwrap();
loop {
let s = sh.seq.0.load(Ordering::Relaxed);
sh.seq.0.store(s + 1, Ordering::Relaxed);
fence(Ordering::Release);
sh.data.0.store(s + 2, Ordering::Relaxed);
sh.seq.0.store(s + 2, Ordering::Release);
/*unsafe {
core::arch::asm!("dc cvau, {}", in(reg) &sh.data);
}*/
}
i tried adding a dc cvau to try making it hit more but it didnt help
so, does it work now?
no
it loads a value of data higher than both seq1 and seq2
or lower if i mess with the dc cvau
Thanks
Why would you even need fences here, just use acquire directly
Well and release
because thats the point of what im trying to show
acquire would fix it
lol
the test is still incorrect
and doesn't match my pseudocode
because seq1 == seq2 is not the right test ;D
Pitust do u know what seqlock does 
you're missing if seq1 & 1 { continue }
okay yea works now
It makes concurrent reads and writes to data larger than a pointer possible by making readers retry
If seqnum is odd or changed from the initial
see, my code was correct :^)
What does your code look like now
std::thread::Builder::new()
.spawn(move || loop {
let seq1 = sh.seq.0.load(Ordering::Acquire);
if seq1 & 1 != 0 {
continue;
}
let data = sh.data.0.load(Ordering::Relaxed);
fence(Ordering::Acquire);
dmb();
let seq2 = sh.seq.0.load(Ordering::Relaxed);
if seq1 == seq2 {
assert!(data == seq1, "loaded {seq1} vs {data} vs {seq2}");
}
})
.unwrap();
loop {
let s = sh.seq.0.load(Ordering::Relaxed);
sh.seq.0.store(s + 1, Ordering::Relaxed);
fence(Ordering::Release);
sh.data.0.store(s + 2, Ordering::Relaxed);
sh.seq.0.store(s + 2, Ordering::Release);
}
