scx-upstream

mirror of https://github.com/sched-ext/scx.git synced 2024-11-25 20:20:23 +00:00

Author	SHA1	Message	Date
Tejun Heo	e26fba9255	Sync from kernel (73f4013eb1eb) This pulls in the support for dump ops.	2024-05-17 01:57:36 -10:00
Andrea Righi	42cee1c2dd	Merge pull request #286 from sched-ext/rustland-low-power-mode scx_rustland: introduce low power mode	2024-05-16 08:28:32 +02:00
Changwoo Min	0ea0a48dfc	Merge pull request #290 from vax-r/Redundant_substract Avoid redundant substraction in rsigmoid_u64	2024-05-16 14:24:12 +09:00
I Hsin Cheng	6cce01c66b	Avoid redundant substraction in rsigmoid_u64 Originally the implementation of function rsigmoid_u64 will perform substraction even when the value of "v" equals to the value of "max" , in which the result is certainly zero. We can avoid this redundant substration by changing the condition from ">" to ">=" since we know when the value of "v" and "max" are equal we can return 0 without any substract operation.	2024-05-16 11:58:39 +08:00
Tejun Heo	147ab7436f	Merge pull request #287 from sirlucjan/journald-rework journal.conf: increase the size of the logs and drop unneeded options	2024-05-15 11:11:22 -10:00
David Vernet	dd724ba08d	Merge pull request #288 from jfernandez/process_runqlat scripts: whitespace cleanup for process_runqlat.bt	2024-05-15 16:00:51 -05:00
Jose Fernandez	7b4689cbde	scripts: whitespace cleanup for process_runqlat.bt Use tabs instead of spaces for indentation in process_runqlat.bt. I also updated comments and fixed a typo. Signed-off-by: Jose Fernandez <josef@netflix.com>	2024-05-15 14:59:13 -06:00
Piotr Gorski	6c4101f30b	journal.conf: increase the size of the logs and drop unneeded options Signed-off-by: Piotr Gorski <lucjan.lucjanov@gmail.com>	2024-05-15 21:35:13 +02:00
Tejun Heo	ad39dd0851	Merge pull request #285 from jfernandez/process_runqlat scripts: Add script to measure runqlat for a process	2024-05-15 09:24:47 -10:00
Andrea Righi	e9ac6105c7	scx_rustland_core: introduce low-power mode Introduce a low-power mode to force the scheduler to operate in a very non-work conserving way, causing a significant saving in terms of power consumption, while still providing a good level of responsiveness in the system. This option can be enabled in scx_rustland via the --low_power / -l option. The idea is to not immediately re-kick a CPU when it enters an idle state, but do that only if there are no other tasks running in the system. In this way, latency-critical tasks can be still dispatched immediately on the other active CPUs, while CPU-bound tasks will be forced to spend more time waiting to be scheduled, basically enforcing a special CPU throttling mechanism that affects only the tasks that are not latency critical. The consequence is a reduction in the overall system throughput, but also a significant reduction of power consumption, that can be useful for mobile / battery-powered devices. Test case (using `scx_rustland -l`): - play a video game (Terraria) while recompiling the kernel - measure game performance (fps) and core power consumption (W) - compare the result of normal mode vs low-power mode Result: Game performance \| Power consumption \| ------------+-----------------+-------------------+ normal mode \| 60 fps \| 6W \| low-power mode \| 60 fps \| 3W \| As we can see from the result the reduction of power consumption is quite significant (50%), while the responsiveness of the game (fps) remains the same, that means battery life can be potentially doubled without significantly affecting system responsiveness. The overall throughput of the system is, of course, affected in a negative way (kernel build is approximately 50% slower during this test), but the goal here is to save power while still maintaining a good level of responsiveness in the system. For this reason the low-power mode should be considered only in emergency conditions, for example when the system is close to completely run out of power or simply to extend the battery life of a mobile device without compromising its responsiveness. Signed-off-by: Andrea Righi <andrea.righi@canonical.com>	2024-05-15 20:32:05 +02:00
Andrea Righi	45a9b178cd	scx_rustland_core: expose nr_running metric Expose the new metric nr_running to keep track of the amount of currently running tasks. Signed-off-by: Andrea Righi <andrea.righi@canonical.com>	2024-05-15 20:28:10 +02:00
Tejun Heo	0d4f6829a8	Merge pull request #284 from vax-r/Fix_typo Fix typo	2024-05-15 06:47:54 -10:00
Jose Fernandez	1beb4ed205	scripts: Add script to measure runqlat for a process Add a `scripts` folder to hold scripts that are useful for the project, and add a bpftrace program to measure the runqueue latency for a process. bpftrace's built-in runqlat.bt program instruments runqueue latency for all processes and does not provide a way to filter by PID. For sched_ext performance work, we are interested in the runqueue latency of a specific process we are trying to optimize, such as a video game. Therefore, we need to create a custom bpftrace program to achieve this. `process_runqlat.bt` instruments runqueue latency for a PID and its threads. This script measures the runqueue latency for a specified PID and includes all threads spawned by that process. USAGE: sudo ./scripts/process_runqlat.bt <PID> The program output will include: - Stats by thread (count, avg latency, total latency) - A histogram of all latency measurements - Aggregated total stats (count, avg latency, total latency) Example output when targeting Terraria's main process: $ sudo ./scripts/process_runqlat.bt 652644 Attaching 5 probes... Instrumenting runqueue latency for PID 652644. Hit Ctrl-C to end. @tasks[AsyncActionDisp, 652676]: count 24, average 2, total 67 @tasks[Finalizer, 652646]: count 24, average 6, total 151 @tasks[Main Thread, 652668]: count 1432, average 8, total 12561 @tasks[Terraria.b:gl0, 652672]: count 1421, average 9, total 13120 @tasks[Main Thread, 652667]: count 1037, average 9, total 10091 @tasks[FACT Thread, 652679]: count 1033, average 10, total 11005 @tasks[Terraria:gdrv0, 652671]: count 3511, average 10, total 37047 @tasks[SDLAudioP3, 652678]: count 982, average 10, total 10210 @tasks[Main Thread, 652666]: count 104, average 10, total 1088 @tasks[SDLAudioP2, 652675]: count 982, average 10, total 10461 @tasks[Main Thread, 652644]: count 5840, average 11, total 69177 @tasks[Main Thread, 694917]: count 44, average 14, total 659 @tasks[Terraria.b:cs0, 652650]: count 3288, average 14, total 47300 @tasks[Thread Pool Wor, 695239]: count 1001, average 35, total 35873 @tasks[Thread Pool Wor, 696059]: count 986, average 35, total 34848 @tasks[Thread Pool Wor, 695458]: count 985, average 36, total 35836 @tasks[Thread Pool Wor, 696058]: count 982, average 37, total 36518 @usec_hist: [0] 16 \| \| [1] 360 \|@ \| [2, 4) 4650 \|@@@@@@@@@@@@@@ \| [4, 8) 14517 \|@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@ \| [8, 16) 16593 \|@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@@\| [16, 32) 5056 \|@@@@@@@@@@@@@@@ \| [32, 64) 6795 \|@@@@@@@@@@@@@@@@@@@@@ \| [64, 128) 7106 \|@@@@@@@@@@@@@@@@@@@@@@ \| [128, 256) 1926 \|@@@@@@ \| [256, 512) 130 \| \| [512, 1K) 2 \| \| @usec_total_stats: count 57152, average 29, total 1699656 Signed-off-by: Jose Fernandez <josef@netflix.com>	2024-05-15 09:18:29 -06:00
vax-r	f293995b59	Fix typo Fix the usage of "scheduler" in the comment of main.bpf.c , it should a verb which is "schedule".	2024-05-15 23:02:35 +08:00
Changwoo Min	971ea6629e	Merge pull request #283 from multics69/scx-lavd-misc scx_lavd: add non-functional misc updates	2024-05-15 17:18:41 +09:00
Changwoo Min	08e7e23cbe	scx_lavd: priint out the current limitaiton of scx_lavd for users Signed-off-by: Changwoo Min <changwoo@igalia.com>	2024-05-15 12:04:09 +09:00
Changwoo Min	a4560c7f7f	scx_lavd: add comments describing the idea of preemption Signed-off-by: Changwoo Min <changwoo@igalia.com>	2024-05-15 12:04:03 +09:00
Tejun Heo	5ad79952b7	Merge pull request #282 from sirlucjan/pacman-hooks-systemd Add pacman hooks for systemd	2024-05-13 05:36:42 -10:00
Piotr Gorski	1fbf4f4f9b	Add pacman hooks for systemd Signed-off-by: Piotr Gorski <lucjan.lucjanov@gmail.com>	2024-05-13 15:52:06 +02:00
Andrea Righi	fa1c146cad	Merge pull request #281 from sched-ext/rustland-fix-offline-cpus scx_rustland: properly support offline CPUs	2024-05-12 09:30:44 +02:00
Andrea Righi	2a7b1cc3c4	scx_rustland: properly support offline CPUs During the initialization phase the scheduler needs to be aware of all the available CPUs in the system (also those that are offline), in order to create a proper per-CPU DSQ for all of them. Otherwise, if some cores are offline, we may get errors like the following: swapper/7[0] triggered exit kind 1024: runtime error (invalid DSQ ID 0x0000000000000007) Backtrace: scx_bpf_consume+0xaa/0xd0 bpf_prog_42ff1b9d1ac5b184_rustland_dispatch+0x12b/0x187 Change the code to configure the BpfScheduler object with the total amount of CPUs available in the system and prevent such failure. This fixes #280. Signed-off-by: Andrea Righi <andrea.righi@canonical.com>	2024-05-12 08:42:46 +02:00
Andrea Righi	5de8ff5bd8	Merge pull request #279 from sched-ext/rustland-max-cpu-util scx_rustland: maximize CPU utilization	2024-05-11 17:08:27 +02:00
Andrea Righi	a31bcc6847	scx_rustland: maximize CPU utilization Always dispatch at least one task, even if all the CPUs are busy. This small overcommitment allows to maximize the CPU utilization without introducing bubbles in the scheduling and also without introducing regressions in terms of resposiveness. Before this change the average CPU utilization of a `stress-ng -c 8` on an 8-cores system is around 95%. With this change applied the CPU utilization goes up to a consistent 100%. Signed-off-by: Andrea Righi <andrea.righi@canonical.com>	2024-05-11 16:23:12 +02:00
Andrea Righi	5f9ce3bba6	Merge pull request #272 from sched-ext/rustland-reduce-scheduling-overhead rustland: reduce scheduling overhead	2024-05-11 10:17:01 +02:00
Andrea Righi	209c454149	scx_rustland_core: fix update_idle description The comment that describes rustland_update_idle() is still incorrectly reporting an old implemention detail. Update its description for better clarity. Signed-off-by: Andrea Righi <andrea.righi@canonical.com>	2024-05-11 07:37:33 +02:00
Andrea Righi	311b7f861c	scx_rustland_core: refine built-in CPU idle selection logic Change the BPF CPU selection logic as following: - if the previously used CPU is idle, keep using it - if the task is not coming from a wait state, try to stick as much as possible to the same CPU (for better cache usage) - if the task is waking up from a wait state rely on the sched_ext built-int idle selection logic This logic can be completely disabled when the full user-space mode is enabled. In this case tasks will always be assigned to the previously used CPU and the user-space scheduler should take care of distributing them among the available CPUs. Signed-off-by: Andrea Righi <andrea.righi@canonical.com>	2024-05-11 07:37:31 +02:00
Tejun Heo	291e2cc996	Merge pull request #278 from sched-ext/topology_numa topology: Support CONFIG_NUMA=n in Topology crate	2024-05-10 12:47:56 -10:00
David Vernet	de512b6bfd	Merge pull request #277 from ptr1337/cachyos-debug INSTALL.md: Add info about debug kernel on arch based	2024-05-10 16:52:22 -05:00
David Vernet	904a89117c	topology: Support CONFIG_NUMA=n in Topology crate Some users are running with NUMA disabled, which makes sense given that it's useless in a lot of contexts. Let's make the Topology crate assume a default node with ID 0 in such cases. Signed-off-by: David Vernet <void@manifault.com>	2024-05-10 16:46:15 -05:00
Peter Jung	c10fcf47f7	INSTALL.md: Add info about debug kernel on arch based Signed-off-by: Peter Jung <admin@ptr1337.dev>	2024-05-10 21:26:13 +02:00
Andrea Righi	63feba9c2b	topology: TopologyMap: add nr_cpus_online() Add a method to TopologyMap to get the amount of online CPUs. Considering that most of the schedulers are not handling CPU hotplugging it can be useful to expose also this metric in addition to the amount of available CPUs in the system. Signed-off-by: Andrea Righi <andrea.righi@canonical.com>	2024-05-10 17:24:20 +02:00
Andrea Righi	f052493005	scx_rustland_core: implement effective time slice on a per-task basis Drop the global effective time-slice and use the more fine-grained per-task time-slice to implement the dynamic time-slice capability. This allows to reduce the scheduler's overhead (dropping the global time slice volatile variable shared between user-space and BPF) and it provides a more fine-grained control on the per-task time slice. Signed-off-by: Andrea Righi <andrea.righi@canonical.com>	2024-05-10 17:24:20 +02:00
Andrea Righi	382ef72999	Merge pull request #276 from sched-ext/scx-utils-sched-running-error scx_utils: report an explicit error when another scheduler is running	2024-05-10 16:17:36 +02:00
Andrea Righi	887812197c	scx_utils: report an explicit error when another scheduler is running If another scheduler is already running, the Rust schedulers based on scx_utils are reporting an error like the following, that can be a bit difficult to understand: Error: Failed to attach struct ops Caused by: bpf call "libbpf_rs::map::Map::attach_struct_ops::{{closure}}" returned NULL Change the scx_ops_attach macro to check if another sched_ext scheduler is running and in that case report a more explicit error. With this applied: $ sudo scx_rustland Error: another sched_ext scheduler is already running Signed-off-by: Andrea Righi <andrea.righi@canonical.com>	2024-05-10 11:20:19 +02:00
Changwoo Min	01faf9408b	Merge pull request #274 from multics69/scx-lavd-preemption02 scx_lavd: support yield-based preemption	2024-05-10 11:32:29 +09:00
Changwoo Min	446de3ef3c	scdx_lavd: minor style changes Signed-off-by: Changwoo Min <changwoo@igalia.com>	2024-05-10 11:07:32 +09:00
Tejun Heo	6ae1031acd	Merge pull request #273 from ptr1337/systemd-service-restart systemd-service: Don't restart always	2024-05-09 06:50:59 -10:00
Changwoo Min	7fcc6e4576	scx_lavd: support yield-based preemption If there is a higher priority task when running ops.tick(), ops.select_cpu(), and ops.enqueue() callbacks, the current running tasks yields its CPU by shrinking time slice to zero and a higher priority task can run on the current CPU. As low-cost, fine-grained preemption becomes available, default parameters are adjusted as follows: - Raise the bar for remote CPU preemption to avoid IPIs. - Increase the maximum time slice. - Gradually enforce the fair use of CPU time (i.e., ineligible duration) Lastly, using CAS, we ensure that a remote CPU is preempted by only one CPU. This removes unnecessary remote preemptions (and IPIs). Signed-off-by: Changwoo Min <changwoo@igalia.com>	2024-05-10 00:54:41 +09:00
Peter Jung	cb8928260e	systemd-service: Don't restart always Currently if the scx.service is failing to launch due issues, systemd will try to start the scheduler all the time. This results into a massive flood to the kernel and does not bring the service up again. explanation of the changes: The StartLimitBurst=2 and StartLimitIntervalSec=30 settings tell systemd that if the service unsuccessfully tries to restart itself twice within 30 seconds, it should enter a failed state and no longer try to restart. This ensures that if the service is truly broken, systemd won't continuously try to restart it. Signed-off-by: Peter Jung <admin@ptr1337.dev>	2024-05-09 14:54:07 +02:00
Andrea Righi	7bc62d8db8	Merge pull request #270 from sched-ext/rustland-user-ringbuffer scx_rustland_core: use a BPF_MAP_TYPE_USER_RINGBUF to dispatch tasks	2024-05-09 06:50:19 +02:00
Andrea Righi	eb2b0b0fa3	Merge pull request #271 from vax-r/Fix_typo Fix typo	2024-05-09 06:49:59 +02:00
vax-r	093a08356e	Fix typo Fix "expermentation" to "experimentation".	2024-05-09 12:10:55 +08:00
Andrea Righi	5da4602ad7	scx_rustland_core: use a BPF_MAP_TYPE_USER_RINGBUF to dispatch tasks Replace the BPF_MAP_TYPE_QUEUE with a BPF_MAP_TYPE_USER_RINGBUF to store the tasks dispatched from the user-space scheduler to the BPF component. This eliminates the need of the bpf() syscalls, significantly reducing the overhead of the user-space->kernel communication and delivering a notable performance boost in the overall system throughput. Based on experimental results, this change allows to reduces the scheduling overhead by approximately 30-35% when the system is overcommitted. This improvement has the potential to make user-space schedulers based on scx_rustland_core viable options for real production systems. Link: https://github.com/libbpf/libbpf-rs/pull/776 Signed-off-by: Andrea Righi <andrea.righi@canonical.com>	2024-05-08 22:16:53 +02:00
David Vernet	07b521b3d0	Merge pull request #266 from sched-ext/rusty_hot_plug rusty: Support CPU hotplug onlining	2024-05-04 21:56:51 -05:00
David Vernet	b9b9875aa7	rusty: Remove task offline tracking scx_rusty's intention is to support hotplug by automatically restarting whenever a hotplug event is encountered. Now that we're not trying to consume a bogus DSQ in the rusty_dispatch() on a newly hotplugged CPU, let's just remove offline tracking. It's really just there as a sanity check, but it triggers if an offline task is made runnable during a hotplug event before the ops.hotplug() callback has been invoked. Signed-off-by: David Vernet <void@manifault.com>	2024-05-04 21:33:55 -05:00
David Vernet	6f1dc6067a	rusty: Check for offline CPU in rusty_dispatch() There's currently a slight issue on existing kernels on the hotplug path wherein we can start to receive scheduling callbacks on a CPU before that CPU has received hotplug events. For CPUs going online, this can possibly confuse a scheduler because it may not be expecting anything to ever happen on that CPU, and therefore may do things that could cause the scheduler to crash. For example, without this patch in scx_rusty, we try to consume from a bogus DSQ that doesn't exist, which causes ext.c to boot out the scheduler. Though this issue will soon be fixed in ext.c, let's explicitly avoid dispatching from an onlining CPU in rusty so that we properly support hotplug on older kernels as well. Signed-off-by: David Vernet <void@manifault.com>	2024-05-04 21:33:54 -05:00
David Vernet	0d6b00238f	common: Add likely/unlikely macros We can hint to the compiler about paths we'll take in a scheduler. This is a common pattern, so lets provide convenience macros. Signed-off-by: David Vernet <void@manifault.com>	2024-05-04 21:33:53 -05:00
David Vernet	4b16f5117a	rusty: Fix alignment Found a misaligned conditional in main.rs. Fix it. Signed-off-by: David Vernet <void@manifault.com>	2024-05-04 21:33:19 -05:00
Changwoo Min	01e5a46371	Merge pull request #263 from multics69/scx_lavd-power01 scx_lavd: support CPU frequency scaling	2024-05-05 10:16:00 +09:00
Tejun Heo	b0a759b40d	Merge pull request #268 from ptr1337/readme-nix README: Add missing link to Nix Install instructions	2024-05-04 04:34:14 -10:00

1 2 3 4 5 ...

760 Commits