scx-upstream

mirror of https://github.com/sched-ext/scx.git synced 2024-12-12 19:47:18 +00:00

Author	SHA1	Message	Date
Andrea Righi	c3cab45f6a	scx_rustland_core: bump up version to 2.0.1 Bump up scx_rustland_core version to include this critical fix that allows to prevent scheduler stalls: `94a3594` ("scx_rustland_core: always dispatch per-cpu kthreads directly") Signed-off-by: Andrea Righi <andrea.righi@linux.dev>	2024-09-04 08:00:25 +02:00
Andrea Righi	94a359434f	scx_rustland_core: always dispatch per-cpu kthreads directly Do not send per-CPU kthreads to the user-space scheduler, but always dispatch them directly from BPF. In specific environments, sending critical per-CPU kthreads to the user-space scheduler can lead to potential stalls. This occurs because the user-space scheduler might be blocked by an action that these per-CPU kthreads need to perform, but they cannot complete their action if the scheduler needs to schedule them, hence the deadlock. To prevent this deadlock, always dispatch the per-CPU kthreads directly from the BPF component, ensuring that the user-space scheduler does not get blocked by these events. Fixes: `c0a2cfb` ("scx_rustland_core: always schedule per-CPU kthreads to user-space") Signed-off-by: Andrea Righi <andrea.righi@linux.dev>	2024-09-04 07:56:58 +02:00
Andrea Righi	0aa71c832b	scx_rustland_core: bump up major version to 2.0.0 The scx_rustland_core API has been redesigned recently, breaking the compatibility with the past. Considering that Rust crates should update their major version when the previous API becomes incompatible [1], bump up the version to 2.0.0. [1] https://semver.org/ Signed-off-by: Andrea Righi <andrea.righi@linux.dev>	2024-08-31 23:23:26 +02:00
Andrea Righi	0bdbb255bb	scx_rustland_core: update README.md with a FIFO example Include the FIFO example directly in the README.md, instead of linking scx_rlfifo. Including the example directly in the README can be more useful and practical in those cases where internet access is not available or when we need to distribute a more "standalone" documentation. Signed-off-by: Andrea Righi <andrea.righi@linux.dev>	2024-08-26 17:42:51 +02:00
Andrea Righi	820fc5a8b4	scx_rustland_core: allow to propagate vtime to BPF Introduce a vtime attribute to struct DispatchedTask that can be set by the user-space scheduler and it'll be use by the BPF component to dispatch the task via scx_bpf_dispatch_vtime(). In this way a user-space scheduler can decide to apply its own internal task ordering or rely on the BPF vtime priority DSQs (or both). Signed-off-by: Andrea Righi <andrea.righi@linux.dev>	2024-08-26 15:56:05 +02:00
Andrea Righi	2ee07bb1fb	scx_rustland_core: temporarily drop RL_PREEMPT_CPU Temporarily drop the RL_PREEMPT_CPU flag, we need a better way to implement preemption in scx_rustland_core and it's not very effective at the moment, so simply it drop it for now (it'll be re-added later in the future in a proper way). This change does not affect any scheduler, since RL_PREEMPT_CPU is currently unused. Signed-off-by: Andrea Righi <andrea.righi@linux.dev>	2024-08-26 15:54:32 +02:00
Tejun Heo	ca13e13ad6	Merge pull request #559 from sched-ext/htejun/cargo-workspace build: Use workspace to group rust sub-projects	2024-08-25 06:26:18 -10:00
Tejun Heo	43950c65bd	build: Use workspace to group rust sub-projects meson build script was building each rust sub-project under rust/ and scheds/rust/ separately. This means that each rust project is built independently which leads to a couple problems - 1. There are a lot of shared dependencies but they have to be built over and over again for each proejct. 2. Concurrency management becomes sad - we either have to unleash multiple cargo builds at the same time possibly thrashing the system or build one by one. We've been trying to solve this from meson side in vain. Thankfully, in issue #546, @vimproved suggested using cargo workspace which makes the sub-projects share the same target directory and built together by the same cargo instance while still allowing each project to behave independently for development and publishing purposes. Make the following changes: - Create two cargo workspaces - one under rust/, the other under scheds/rust/. Each contains all rust projects underneath it. - Don't let meson descend into rust/. These are libraries used by the rust schedulers. No need to build them from meson. Cargo will build them as needed. - Change the rust_scheds build target to invoke `cargo build` in scheds/rust/ and let cargo do its thing. - Remove per-scheduler meson.build files and instead generate custom_targets in scheds/rust/meson.build which invokes `cargo build -p $SCHED`. - This changes rust binary directory. Update README and meson-scripts/install_rust_user_scheds accordingly. - Remove per-scheduler Cargo.lock as scheds/rust/Cargo.lock is shared by all schedulers now. - Unify .gitignore handling. The followings are build times on Ryzen 3975W: Before: ________________________________________________________ Executed in 165.93 secs fish external usr time 40.55 mins 2.71 millis 40.55 mins sys time 3.34 mins 36.40 millis 3.34 mins After: ________________________________________________________ Executed in 36.04 secs fish external usr time 336.42 secs 0.00 millis 336.42 secs sys time 36.65 secs 43.95 millis 36.61 secs Wallclock time is reduced 5x and CPU time 7x.	2024-08-25 00:47:58 -10:00
Andrea Righi	41dfa2481b	scx_rustland_core: update README.md Signed-off-by: Andrea Righi <andrea.righi@linux.dev>	2024-08-25 12:39:08 +02:00
Andrea Righi	894f9582d0	scx_rustland_core: hide shutdown boilerplate in BpfScheduler Refactor the code to hide the shutdown handling inside BpfScheduler and simply use the exited() method to check when the scheduler is stopped. Signed-off-by: Andrea Righi <andrea.righi@linux.dev>	2024-08-25 12:17:04 +02:00
Andrea Righi	a2e97fecbb	scx_rustland_core: merge verbose and debug in the same option There is no reason to have two separate options for "verbose" and "debug" mode. Just merge the two and always use "debug". If enabled, increase verbosity to stdout and enable reporting BPF scheduling events in debugfs (e.g., /sys/kernel/debug/tracing/trace_pipe). Signed-off-by: Andrea Righi <andrea.righi@linux.dev>	2024-08-25 09:45:20 +02:00
Andrea Righi	cb16a11342	scx_rustland_core: get rid of the global scheduler's slice_us Since scx_rustland_core enables setting a time slice on a per-task basis during task dispatch, there's no need to maintain a global time slice in the BPF component. Instead, a global time slice can simply be managed in user-space, achieving the same outcome. Therefore, drop the global slice_us property from BpfScheduler to simplify the API. NOTE: if a time slice is not specified for a task, SCX_SLICE_DFL will be used by default. Signed-off-by: Andrea Righi <andrea.righi@linux.dev>	2024-08-25 09:45:18 +02:00
Andrea Righi	0aa23481de	scx_rustland_core: drop update_tasks() and introduce notify_complete() The update_tasks() API is somewhat confusing, so replace it with a clearer API, notify_complete(). This new API will return control to the BPF component and inform it about the number of tasks still pending in the user-space scheduler. Signed-off-by: Andrea Righi <andrea.righi@linux.dev>	2024-08-25 00:45:23 +02:00
Andrea Righi	cef8ff8757	scx_rustland_core: get rid of the low_power API The low-power API is a bit of a hack implemented purely in the BPF layer, this should be better re-implemented with some concepts of topology awareness. Therefore, get rid of this API for now. Signed-off-by: Andrea Righi <andrea.righi@linux.dev>	2024-08-24 21:29:10 +02:00
Andrea Righi	568e292a24	scx_rustland_core: get rid of the exiting task API The current API used to notify the user-space scheduler when a task exits is really confusing (setting a negative value in queued_task_ctx.cpu), and it's also possible to detect task exiting events from user-space (or check in procfs, even if it's slower). In any case, a better API should be provided for this, so drop the current one for now. NOTE: this will cause additional memory usage for scx_rustland, but it can be fixed/addressed later in a separate commit (i.e., providing a periodic garbage collector for the unused task entries). Signed-off-by: Andrea Righi <andrea.righi@linux.dev>	2024-08-24 21:29:10 +02:00
Andrea Righi	eec395f16a	scx_rustland_core: better time slice control Instead of determining the task time slice in ops.enqueue(), refresh the time slice immediately before the task is started on its assigned CPU in ops.running(). This ensures to apply the exact time slice specified by the user-space scheduler and the sched_ext core will never implicitly dispatch tasks using SCX_SLICE_DFL. Signed-off-by: Andrea Righi <andrea.righi@linux.dev>	2024-08-24 21:29:10 +02:00
Andrea Righi	c0a2cfb481	scx_rustland_core: always schedule per-CPU kthreads to user-space Signed-off-by: Andrea Righi <andrea.righi@linux.dev>	2024-08-24 21:29:10 +02:00
Andrea Righi	5d544ea264	scx_rustland_core: move CPU idle selection logic in user-space Allow user-space scheduler to pick an idle CPU via self.bpf.select_cpu(pid, prev_task, flags), mimicking the BPF's select_cpu() iterface. Also remove the full_user option and always rely on the idle selection logic from user-space. Signed-off-by: Andrea Righi <andrea.righi@linux.dev>	2024-08-24 21:28:13 +02:00
Tejun Heo	4d1f0639d8	Version: v1.0.3	2024-08-21 06:42:11 -10:00
Tejun Heo	45f7fd13b7	versions: Synchronize crate dependency versions	2024-08-08 14:45:46 -10:00
Tejun Heo	63c4a0191f	Merge branch 'main' into topic/inlined-skeleton-members	2024-08-08 14:23:37 -10:00
Tejun Heo	cd6a4d72c7	Bump versions for 1.0.2 release	2024-08-08 14:10:16 -10:00
Tejun Heo	7c3ffe96e1	Unify crate dependency versions Different sub-projects are using different versions for the same crates. Synchronize them to the latest.	2024-08-08 13:26:47 -10:00
Andrea Righi	51cfb69199	scx_rustland_core: re-introduce partial mode Re-add the partial mode option that was dropped during the refactoring. The partial option allows to apply the scheduler only to the tasks which have their scheduling policy set to SCHED_EXT via sched_setscheduler(). Signed-off-by: Andrea Righi <andrea.righi@linux.dev>	2024-08-07 08:41:06 +02:00
Andrea Righi	e1f2b3822e	scx_rustland_core: drop CPU ownership API The API for determining which PID is running on a specific CPU is racy and is unnecessary since this information can be obtained from user space. Additionally, it's not reliable for identifying idle CPUs. Therefore, it's better to remove this API and, in the future, provide a cpumask alternative that can export the idle state of the CPUs to user space. As a consequence also change scx_rustland to dispatch one task a time, instead of dispatching tasks in batches of idle cores (that are usually not accurate due to the racy nature of the CPU ownership interaface). Dispatching one task at a time even makes the scheduler more performant, due to the vruntime scheduling being applied to more tasks sitting in the scheduler's queue. Signed-off-by: Andrea Righi <andrea.righi@linux.dev>	2024-08-07 08:41:06 +02:00
Andrea Righi	9a0e7755df	scx_rustland_core: export counter of online CPUs Introduce a helper to get the amount of online CPUs tracked by the BPF part. Signed-off-by: Andrea Righi <andrea.righi@linux.dev>	2024-08-07 08:10:53 +02:00
Andrea Righi	d9c9f78e3e	scx_rustland: re-align vruntime and time slice evaluation to scx_bpfland Drop the slice boost logic and apply a vruntime and task time slice evaluation approach similar to scx_bpfland (but implement this in the user-space component instead of the BPF part). Additionally, introduce a slice_us_min parameter to define the minimum time slice that can be assigned to a task, also similar to scx_bpfland. Signed-off-by: Andrea Righi <andrea.righi@linux.dev>	2024-08-07 08:10:53 +02:00
Andrea Righi	e1e6e31208	scx_rustland_core: update copyright info Signed-off-by: Andrea Righi <andrea.righi@linux.dev>	2024-08-07 08:10:53 +02:00
Andrea Righi	b87541a26e	scx_rustland_core: refactor idle CPU selection logic Use the same idle selection logic used in scx_bpfland also in scx_rustland_core. Also drop fifo_mode and always use the BPF idle selection logic by default as long as the system is not saturated, unless full_user is specified. This approach allows user-space schedulers aiming for maximum performance to leverage the BPF idle selection logic (bypassing user-space), while those seeking full control can enable full_user to bypass the BPF CPU idle selection logic and choose the target CPU for each task from user-space. Signed-off-by: Andrea Righi <andrea.righi@linux.dev>	2024-08-07 08:10:53 +02:00
Andrea Righi	d8985306f4	scx_rustland: user-space interactive task classifier We don't need to send the number of voluntary context switches (nvcsw) from BPF to user-space, as this information is already accessible in user-space via procfs. Sending this data would only create unnecessary overhead for schedulers that don't require it, and those that do can easily retrieve it through procfs. Therefore, drop this metric from scx_rustland_core and change scx_rustland implementing an interactive task classifier fully in the user-space part of the scheduler. Also drop some options that are not provide any significant benefit (also in preparation of a bigger refactoring to define a better API for the user-space framework). Signed-off-by: Andrea Righi <andrea.righi@linux.dev>	2024-08-06 17:56:58 +02:00
Andrea Righi	a8d14fc0c4	Merge pull request #471 from sched-ext/rustland-core-musl scx_rustland_core: add support for musl	2024-08-06 17:56:19 +02:00
Kawanaao	c3109ebeed	scx_rustland_core alloc: Replaced RefCell with Mutex Necessary for some multi-threaded cases	2024-08-06 13:10:21 +00:00
Andrea Righi	4c7fb5cdbd	scx_rustland_core: add support for musl It seems that musl glibc is not POSIX.1-2001, POSIX.1-2008 compliant and using sched_setscheduler() just returns -ENOSYS: https://git.musl-libc.org/cgit/musl/commit/src/sched/sched_setscheduler.c?id=1e21e78bf7a5c24c217446d8760be7b7188711c2 Switch to pthread_setschedparam() to properly support building scx_rustland_core with musl. This fixes #469. Signed-off-by: Andrea Righi <andrea.righi@linux.dev>	2024-08-06 07:32:38 +02:00
Andrea Righi	d4005dd186	scx_rustland_core: fix missing import (timespec) Explicitly import timespec to fix the following potential build error: error[E0422]: cannot find struct, variant or union type `timespec` in this scope --> src/bpf.rs:365:35 \| 365 \| sched_ss_repl_period: timespec { \| ^^^^^^^^ not found in this scope \| help: consider importing this struct \| 6 + use libc::timespec; \| This fixes issue #469. Signed-off-by: Andrea Righi <andrea.righi@linux.dev>	2024-08-05 23:15:48 +02:00
Daniel Müller	565aec3662	rust: Update libbpf-rs & libbpf-cargo to 0.24 Update libbpf-rs & libbpf-cargo to 0.24. Among other things, generated skeletons now contain directly accessible map and program objects, no longer necessitating the use of accessor methods. As a result, the risk for mutability conflicts is reduced greatly. Signed-off-by: Daniel Müller <deso@posteo.net>	2024-07-16 11:48:52 -07:00
Tejun Heo	51334b5c4d	Bump versions for 1.0.1 release	2024-07-15 13:21:52 -10:00
Tejun Heo	761ec142ce	Bump most versions to 1.0.0 sched_ext is about to be merged upstream. There are some compatibility breaking changes and we're making the current sched_ext/for-6.11 1edab907b57d ("sched_ext/scx_qmap: Pick idle CPU for direct dispatch on !wakeup enqueues") the baseline. Tag everything except scx_mitosis as 1.0.0. As scx_mitosis is still in early development and is currently temporarily disabled, only the patchlevel is bumped.	2024-07-12 11:34:14 -10:00
I Hsin Cheng	1595da78dc	scx_rustland_core: Remove unused variable Remove unused variable "tctx" in rustland_select_cpu. Signed-off-by: I Hsin Cheng <richard120310@gmail.com>	2024-07-05 01:04:49 +08:00
Andrea Righi	5db0908530	scx_rustland_core: make sure to use a valid CPU during direct dispatch We may end up selecting an invalid CPU (according to the task's cpumask) when dispatching the task via dispatch_direct_cpu(). When this happens simply return an error and do not dispatch the task and let the caller handle the error: in the context of select_cpu() we can simply ignore the dispatch and return the target CPU; in the context of FIFO mode dispatch we can fallback to SCX_DSQ_LOCAL if the target CPU is not valid. This fixes issue #353. Signed-off-by: Andrea Righi <andrea.righi@canonical.com>	2024-06-25 14:11:46 +02:00
Andrea Righi	e4b13b2aa6	scx_rustland_core: reduce dispatch overhead Kick CPUs in the dispatch path only when needed (typically when tasks are bounced to other CPUs). Moreover, avoid to consume all the tasks dispatched at once. This seems to reduce the BPF overhead (according to bpftop), going from ~10% CPU usage down to ~6% CPU usage of rustland_dispatch() on an over commissioned system, without introducing any measureable performance regression. Tested-by: SoulHarsh007 <harsh.peshwani@outlook.com> Signed-off-by: Andrea Righi <andrea.righi@canonical.com>	2024-06-25 14:02:39 +02:00
Andrea Righi	631d5576dc	scx_rustland_core: refactor CPU selection logic Allow to dispatch tasks directly (bypassing the user-space scheduler) only when the scheduler is operating in FIFO mode. On an over-commissioned system, directly dispatching tasks can only increase OS noise. These tasks can get a brief priority boost and an extended time slice just because they found an idle CPU, which can lead to erratic behavior. This is particularly problematic when measuring performance stability, such as evaluating the frames-per-second (fps) of a video game on an overloaded system. In such cases, it's better to bounce all tasks to the user-space scheduler, that will ensure a better level of fairness and smoother performance. Moreover, get rid of the second chance dispatch logic introduced in commit `4791d862` ("scx_rustland_core: second chance CPU migration"). This seems to provide benefits only on certain architectures (Intel), but it can introduces lags in others (AMD). Tested-by: SoulHarsh007 <harsh.peshwani@outlook.com> Signed-off-by: Andrea Righi <andrea.righi@canonical.com>	2024-06-25 14:02:39 +02:00
Andrea Righi	19217b5722	scx_rustland_core: clarify comment about resuming FIFO mode Tested-by: SoulHarsh007 <harsh.peshwani@outlook.com> Signed-off-by: Andrea Righi <andrea.righi@canonical.com>	2024-06-25 14:02:39 +02:00
Andrea Righi	081d4bdb86	scx_rustland_core: add debugging to dispatch_direct_cpu() Report dispatch_direct_cpu() events in the trace, like any other dispatch-related event. Tested-by: SoulHarsh007 <harsh.peshwani@outlook.com> Signed-off-by: Andrea Righi <andrea.righi@canonical.com>	2024-06-25 14:02:39 +02:00
Andrea Righi	b04e82b5eb	scx_rustland_core: include buddy-alloc and refactor allocator code The dependency of the buddy-alloc crate [1] seems to cause some troubles with packaging, mostly because the selftests for the crate are failing when it's compiled in release mode. For example: $ cargo test --release -- --nocapture thread 'tests::fast_alloc::test_basic_malloc' panicked at src/tests/fast_alloc.rs:25:13: assertion `left == right` failed left: 0 right: 42 Some of these failures with BuddyAlloc can be fixed by using a memory arena buffer aligned to page size. However, some test failures with FastAlloc persist that cannot be resolved merely by aligning the pre-allocated memory arena to the page size, as mentioned in [2]. The concern is that this may potentially lead to actual memory bugs. Therefore, it seems safer to refactor the custom allocator code to simply use BuddyAlloc, dropping FastAlloc completely. To achieve this, the entire BuddyAlloc code has been directly included in scx_rustland_core, referencing the original project and its MIT licensing information (with the entire code still distributed under the GPLv2 license). Then the code has been slightly modified to remove FastAlloc and the external dependency on the buddy-alloc crate has been dropped. From a performance perspective this change doesn't seem to introduce any measurable regression. [1] https://github.com/jjyr/buddy-alloc [2] https://github.com/jjyr/buddy-alloc/issues/16 Signed-off-by: Andrea Righi <andrea.righi@canonical.com>	2024-06-19 14:44:04 +02:00
Tejun Heo	d18bb4831a	Merge pull request #363 from sched-ext/htejun/compat-strip Strip compat support	2024-06-16 20:04:50 -10:00
Andrea Righi	812aeb3b81	scx_rustland_core: fix potential race in dispatch_task() Fix a potential race condition that might lead to a task being dispatched without kicking the target CPU, which could result in a potential stall. With this applied, scx_rustland has been running without any stall for about 18 hours on a system where the issue was previously quite easy to reproduce. Moreover, clarify a couple of comments in the dispatch path. This fixes issue #353. Tested-by: SoulHarsh007 <harsh.peshwani@outlook.com> Signed-off-by: Andrea Righi <andrea.righi@canonical.com>	2024-06-16 08:31:48 +02:00
Tejun Heo	5b5e5be906	compat: Drop __COMPAT_SCX_KICK_IDLE In preparation of upstreaming, let's set the min version requirement at the released v6.9 kernels. Drop __COMPAT_SCX_KICK_IDLE. The open helper macros now check the existence of SCX_KICK_IDLE and abort if not.	2024-06-15 20:24:15 -10:00
Tejun Heo	7c9aedaefe	compat: Drop __COMPAT_scx_bpf_switch_all() In preparation of upstreaming, let's set the min version requirement at the released v6.9 kernels. Drop __COMPAT_scx_bpf_switch_call(). The open helper macros now check the existence of SCX_OPS_SWITCH_PARTIAL and abort if not.	2024-06-15 20:03:37 -10:00
Andrea Righi	6cc2595922	scx_rustland_core: allow user-space scheduler to preempt other tasks The same change was previously applied with commit `8820af8d` ("scx_rustland: enable user-space scheduler to preempt other tasks"). However, it was reverted in commit `732ba490` ("scx_rustland: avoid using SCX_ENQ_PREEMPT") with the introduction of the global dynamic time slice (inversely proportional to the number of running tasks), because it already provided sufficient context switch opportunities, negating any advantage of the preemption. With the introduction of the virtual time slice in commit `6f4cd853` ("scx_rustland: introduce virtual time slice"), re-introducing the ability for the user-space scheduler to preempt other tasks now appears beneficial again. As confirmed by experimental results, this change helps prevent potential audio cracking issues and enhances overall system responsiveness, resulting in a 4-5% increase in fps performing the usual benchmark of gaming while recompiling the kernel. Tested-by: SoulHarsh007 <harsh.peshwani@outlook.com> Signed-off-by: Andrea Righi <andrea.righi@canonical.com>	2024-06-14 20:09:19 +02:00
Andrea Righi	2d2c6083d7	scx_rustland_core: use SCX_DSQ_LOCAL only with kthreads Tasks dispatched directly using SCX_DSQ_LOCAL may receive excessively high priority compared to those dispatched by the user-space scheduler. To avoid this priority disparity, dispatch tasks to the per-CPU DSQs using direct dispatch and reserve SCX_DSQ_LOCAL for per-CPU kthreads only. Tested-by: Tested-by: SoulHarsh007 <harsh.peshwani@outlook.com> Signed-off-by: Andrea Righi <andrea.righi@canonical.com>	2024-06-14 20:09:19 +02:00

1 2 3

101 Commits