scx-upstream

mirror of https://github.com/sched-ext/scx.git synced 2024-11-26 04:30:23 +00:00

Author	SHA1	Message	Date
Tejun Heo	a71a40c581	Merge pull request #33 from davemarchevsky/fewer_kptr_xchg scx_flatcg: Keep cgroup rb nodes stashed	2024-03-13 06:45:41 -10:00
David Vernet	91cb5ce8ab	Merge pull request #178 from sched-ext/multi_numa_rusty rusty: Implement NUMA-aware load balancing	2024-03-12 15:50:27 -05:00
David Vernet	c8d841d50b	rusty: Add comments + use VecDeque Given the complexity of migrating load between nodes (we're doing four nested loops), we should add a comment explaining what we're doing. This commit does that. In addition, we use a VecDeque to store (and then restore) push nodes and push domains so that we can re-add them to their respective lists in load-sorted order rather than reverse-load-sorted order. This allows us to avoid having to do unnecessary right-shifts every time a push object is re-added to its containing list. Signed-off-by: David Vernet <void@manifault.com>	2024-03-12 13:49:14 -07:00
Tejun Heo	a9457a408e	scx_layered: stat reporting updates	2024-03-12 10:48:21 -10:00
Tejun Heo	a642fc873b	scx_layered: Fix stat reporting GSTAT_TASK_CTX_FREE_FAILED should report total while EXCL_* should report delta pct. Fix them.	2024-03-12 10:25:51 -10:00
David Vernet	03f68092ee	rusty: Fix a few remaining issues Fixing alignment, moving a couple bail! calls around, and adding a missing break from move_between_nodes() that lets us bail out of a loop early. Signed-off-by: David Vernet <void@manifault.com>	2024-03-12 12:44:38 -07:00
Tejun Heo	58cbc5361d	scx_layered: warn if omitted stats aren't zero	2024-03-12 09:29:31 -10:00
Tejun Heo	37006d1bc1	scx_layered: Use saturating sub when reading system stats, other misc changes Sometimes io_wait time goes in the wrong direction. Use saturating sub.	2024-03-12 06:14:06 -10:00
Tejun Heo	342a4946af	scx_layered: Better pct formatting when printing stats	2024-03-11 22:18:03 -10:00
Tejun Heo	be2102775b	scx_layered: Implement min_exec_us option which can be used to penalize tasks which wake up very frequently without doing much.	2024-03-11 22:13:11 -10:00
Tejun Heo	0c62b24993	scx_layered: Implement exclusive property A task in an exclusive grouped or open layer occupied a whole core - the sibling CPU is kept idle.	2024-03-11 18:27:16 -10:00
David Vernet	24d798c2ff	rusty: Use a flat list of NumaNodes during LB As Tejun pointed out in review, the disadvantage of using push/pull/balanced lists is that if the domains inside the nodes are balanced, we won't be able to push load between them. I'd originally done it that way both as an optimization, but also to allow me to iterate over the lists of pushable and pullable domains mutably. That was addressed in the prior commit, but the nodes themselves were still put into 3 buckets. I think this is generally just a cleaner way of doing things, so let's just collapse the nodes into a flat list as well. This prevents us from having to coalesce the lists, std::mem::swap them, etc. Signed-off-by: David Vernet <void@manifault.com>	2024-03-11 21:04:10 -07:00
David Vernet	829b1d3ced	rusty: Don't use multiple SortedVec's in struct NumaNode Tejun pointed out that a possible issue exists in the current implementation, wherein if you have two NUMA nodes that are imbalanced, but their domains are internally balanced, we'll fail to migrate between them if all nodes are in the balanced_nodes list. To address this, let's just use a single global list for all types of domains, and do checking internally for imbalances. The reason it was done this way in the first place was to allow me to mutably iterate over both vectors in a nested loop. The way around that is to just use loop {} and push/pop domains from the list. We could do the same thing for the NUMA nodes themselves, which are also in 3 separate lists in the LoadBalancer. We'll do that in a subsequent commit. Signed-off-by: David Vernet <void@manifault.com>	2024-03-11 21:04:10 -07:00
David Vernet	3d2507e6f2	rusty: Add separate flag for x NUMA greedy task stealing In scx_rusty, a CPU that is going to go idle will attempt to steal tasks from remote domains when its domain has no tasks to run, and a remote domain has at least greedy_threshold enqueued tasks. This stealing is temporary, but of course has a cost in that the CPU that's stealing the task may cause it to suffer from cache misses, or in the case of multi-node machines, remote NUMA accesses and working sets split across multiple domains. Given the higher cost of x NUMA work stealing, let's add a separate flag that lets users tune the threshold for doing cross NUMA greedy task stealing. Signed-off-by: David Vernet <void@manifault.com>	2024-03-11 21:02:23 -07:00
Tejun Heo	76cc337d78	scx_layered: Add exclusive option to Open and Grouped layers Actual implementation isn't done yet.	2024-03-11 12:07:03 -10:00
Jordan Rome	54fe1c954e	Merge pull request #179 from jordalgo/bpftool Fetch and build bpftool by default	2024-03-11 17:54:29 -04:00
Tejun Heo	56a83ef07e	Merge pull request #183 from sched-ext/rustland-core-revert-consume-raw Revert "scx_rustland_core: use new consume_raw() libbpf-rs API"	2024-03-11 11:00:48 -10:00
Andrea Righi	bd2c18afd5	Revert "scx_rustland_core: use new consume_raw() libbpf-rs API" In order to use the new consume_raw() API we need to depend on a version of libbpf-rs that is not released yet. Apparently adding such dependency may introduce a potential dependency conflict with libbpf-sys. Therefore, revert this change and go back to the previous consume() API. One a new version of libbpf-rs will be out we can update all our dependencies to use the new libbpf-rs and re-apply this patch to scx_rustland_core. Fixes: `7c8c5fd` ("scx_rustland_core: use new consume_raw() libbpf-rs API") Signed-off-by: Andrea Righi <andrea.righi@canonical.com>	2024-03-11 21:54:21 +01:00
Jordan Rome	ffc7b7dc4a	Fetch and build bpftool by default This pairs with the new default behavior to fetch and build libbpf and is mostly being used so we can use the latest bpftool and libbpf.	2024-03-11 10:00:01 -07:00
Dave Marchevsky	3b7f33ea1b	scx_flatcg: Keep cgroup rb nodes stashed The flatcg scheduler uses a rb_node type - struct cgv_node - to keep track of vtime. On cgroup init, a cgv_node is created and stashed in a hashmap - cgv_node_stash - for later use. In cgroup_enqueued and try_pick_next_cgroup, the node is inserted into the rbtree, which required removing it from the stash before this patch's changes. This patch makes cgv_node refcounted, which allows keeping it in the stash for the entirety of the cgroup's lifetime. Unnecessary bpf_kptr_xchg's and other boilerplate can be removed as a result. Note that in addition to bpf_refcount patches, which have been upstream for quite some time, this change depends on a more recent series [0]. [0]: https://lore.kernel.org/bpf/20231107085639.3016113-1-davemarchevsky@fb.com/ Signed-off-by: Dave Marchevsky <davemarchevsky@fb.com>	2024-03-11 08:35:25 -07:00
Andrea Righi	b7c06b9ed9	Merge pull request #181 from sched-ext/rustland-interactive-tuning scx_rustland: interactive tuning	2024-03-10 19:31:00 +01:00
Andrea Righi	743e3528dc	Merge pull request #180 from sched-ext/rustland-core-new-libbpf-rs-api scx_rustland_core: use new consume_raw() libbpf-rs API	2024-03-10 19:30:40 +01:00
Tejun Heo	b2337d50f4	Merge pull request #182 from sched-ext/htejun-patch-1 Update README.md	2024-03-10 08:27:09 -10:00
Tejun Heo	423cd3ab67	Update README.md Arch now defaults to clang17. Drop the custom clang installation.	2024-03-10 08:26:40 -10:00
Andrea Righi	155444e1c0	scx_rustland: set default time slice to 5ms In line with rustland's focus on prioritizing interactive tasks, set the default base time slice to 5ms. This allows to mitigate potential audio craking issues or system lags when the system is overloaded or under memory pressure condition (i.e., https://github.com/sched-ext/scx/issues/96#issuecomment-1978154324). A downside of this change is to introduce potential regressions in the throughput of CPU-intensive workloads, but in such scenarios rustland may not be the optimal choice and alternative schedulers may be preferred. Signed-off-by: Andrea Righi <andrea.righi@canonical.com>	2024-03-10 14:46:11 +01:00
Andrea Righi	0a7161cbc7	scx_rustland: limit range of task weight Some high-priority tasks may have a weight too high, that can potentially disrupt the slice boost optimization logic, causing interactive tasks to be less responsive. In line with rustland's focus on prioritizing interactive tasks, prevent giving too much CPU bandwidth to such high-priority tasks by limiting the maximum task weight to 1000. This allows to maintain a good level of system responsiveness even in presence of tasks with a really high priority. Signed-off-by: Andrea Righi <andrea.righi@canonical.com>	2024-03-10 14:39:29 +01:00
Andrea Righi	7c8c5fdd48	scx_rustland_core: use new consume_raw() libbpf-rs API Use the new consume_raw() API provided by libbpf-rs with https://github.com/libbpf/libbpf-rs/pull/680. This allows to be more precise and efficient at processing tasks consumed from the BPF ring buffer. NOTE: the new consume_raw() API is not available yet in any official release of the libbpf-rs crate, but cargo allows to pick versions directly from git. This slightly increases the build time of scx_rustland_core and the schedulers based on this crate (since we need to recompile libbpf-rs from source), but we can re-add a proper versioned dependency once the libbpf-rs is out. TODO: this new API also offers the possibility to consume multiple items from the BPF ring buffer with a single call to consume_raw(). This could be investigated and implemented as a potential future enhancement. Signed-off-by: Andrea Righi <andrea.righi@canonical.com>	2024-03-10 09:55:17 +01:00
David Vernet	1c3168d2a4	topology: Don't assume unique core IDs The current topology.rs crate assumes that all cores have unique core IDs in a system. This need not be the case, such as in certain Intel Xeon processors which reuse core IDs in different NUMA nodes. Let's update the crate to assume unique core IDs only per socket. Signed-off-by: David Vernet <void@manifault.com>	2024-03-08 15:13:46 -06:00
David Vernet	26a94b1b14	rusty: Add debug! logging to load_balance.rs We removed the debug!() output that was previously present in main.rs. Let's add more debug!() output that helps debug the current LB hierarchy. Signed-off-by: David Vernet <void@manifault.com>	2024-03-08 15:13:46 -06:00
David Vernet	0d0b101398	rusty: Add load balancing statistics to rusty The scx_rusty load balancer is currently no longer exporting statistics such as domain load averages, load sums, etc. Now that we're also balancing by NUMA, we'll need a way to hierarchically illustrate load balancing statistics. This patch adds support for that. Signed-off-by: David Vernet <void@manifault.com> updating stats printing Signed-off-by: David Vernet <void@manifault.com>	2024-03-08 15:13:36 -06:00
David Vernet	0871a9525d	rusty: Add direct_greedy_numa flag Users may want to toggle whether tasks can be temporarily sent to idle CPUs on remote NUMA nodes. By default, we want it to be disabled as a task spanning multiple NUMA nodes will end up having its working set spanning both nodes, which is probably not desirable. However, in case a workload really wants to encourage work conservation, let's add a flag that allows them to toggle it. Signed-off-by: David Vernet <void@manifault.com>	2024-03-08 15:12:00 -06:00
David Vernet	d0ebfb85ef	rusty: Disable direct greedy stealing between NUMA nodes scx_rusty currently pushes tasks to idle cores if the direct greedy threshold is exceeded, even if the core is on a remote NUMA node. This behavior is probably not desired in most scenarios. The worst that will happen if a task is pushed to an idle core in the same node is some L3 cache miss traffic, but for multiple NUMA nodes, it could cause the task to have its working set span multiple nodes. Let's disable direct greedy work stealing across NUMA nodes. A future commit will add a flag that's disabled by default, and let's users turn this on if they really want to encourage work conservation. Signed-off-by: David Vernet <void@manifault.com>	2024-03-08 15:11:59 -06:00
David Vernet	12e0586fe9	cpumask: Update cpumask fmt function The cpumask print formatter doesn't look great in its current form, which uses the BitVec formatter under the hood: [INFO] NUMA[00 32:<[1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1]>] [INFO] DOM[00] 32:<[1, 1, 1, 1, 1, 1, 1, 1, 0, 0, 0, 0, 0, 0, 0, 0, 1, 1, 1, 1, 1, 1, 1, 1, 0, 0, 0, 0, 0, 0, 0, 0]> [INFO] DOM[01] 32:<[0, 0, 0, 0, 0, 0, 0, 0, 1, 1, 1, 1, 1, 1, 1, 1, 0, 0, 0, 0, 0, 0, 0, 0, 1, 1, 1, 1, 1, 1, 1, 1]> Let's just iterate over the mask and manually format the string using the binary formatter over the slice of u64's, which renders like this: [INFO] NUMA[00] 0b11111111111111111111111111111111] [INFO] DOM[00] 0b00000000111111110000000011111111 [INFO] DOM[01] 0b11111111000000001111111100000000 Signed-off-by: David Vernet <void@manifault.com>	2024-03-08 15:11:17 -06:00
David Vernet	db152cfbe8	rusty: Implement NUMA-aware load balancing Right now, scx_rusty has no notion of domains spanning NUMA nodes, and makes no distinction when making load balancing decisions, or work stealing. This can cause problems on multi-NUMA machines, as load balancing and work stealing across NUMA nodes has significantly different cost from across L3 cache boundaries. In order to better support multi-NUMA machines, this commit adds another layer to the rusty load balancer, which balances across NUMA nodes using a different cost function from balancing across domains. Load balancing now takes place over the span of two passes: 1. In the first pass, we fix imbalances across NUMA nodes by moving tasks between domains across those NUMA node boundaries. We require a load imbalance of at least 17% in order to move load at this stage. The ratio of load imbalance we attempt to adjust (50%) and the maximum amount of load we're allowed to push out of a domain (50%) is still the same as when balancing between domains inside a NUMA node, but this is easy to tune with the current setup. 2. Once we've balanced across NUMA nodes, we iterate over all nodes and balance between the domains within each NUMA node. The cost function here is the same as what it has been thus far: we require at least a 5% imbalance in order to trigger load balancing. There are a few additional changes / improvements to load balancing in this commit: 1. NUMA nodes and domains are now ordered according to their load by using SortedVec objects. We were previously using BTreeMap keyed by load, but this was suboptimal due to the fact that it doesn't allow duplicate entries. 2. We're no longer exporting load balancing statistics as a vector of data such as load sums, averages, and imbalances. This is instead all encapsulated in the load balancing hierarchy we setup in lb.load_balance(). These statistics are not yet exported, but they will be in a subsequent commit. One of the issues with this commit is that it does introduce some almost-identical logic that somehow begs to be deduplicated. For example, when we balance between NUMA nodes, the logic for iterating over push nodes and pushing to pull nodes is very similar to the logic of iterating over push domains and pull domains when balancing within a node. It may be that this can be improved. The following are some benchmarks run on an Intel Xeon Gold 6138 (2 x 40 core processor): kcompile -------- On Commit a27648c74210 ("afs: Fix setting of mtime when creating a file/dir/symlink"): 1. make allyesconfig 2. make -j $(nproc) built-in.a 3. make -j clean 4. goto 2 Runtime ------- o-----------o-----------o----------o \| scx_rusty \| CFS \| Delta \| ---------o-----------o-----------o----------o Mean \| 562.688s \| 566.085s \| -.6% \| ---------o-----------o-----------o----------o Variance \| 0.54387 \| 0.72431 \| -24.9% \| ---------o-----------o-----------o----------o o-----------o-----------o----------o \| rusty NUMA\| rusty ORIG\| Delta \| ---------o-----------o-----------o----------o Mean \| 562.688s \| 563.209s \| -.092% \| ---------o-----------o-----------o----------o Variance \| 0.54387 \| 0.42038 \| 29.38% \| ---------o-----------o-----------o----------o scx_rusty with NUMA awareness clearly beats CFS, but only barely beats scx_rusty without it. This isn't necessarily super surprising given that this is kcompile, which has very poor front-end CPU locality. Further experimentation with toggling the cost function for performing migrations may improve this further. CPU util -------- o-----------o-----------o----------o \| scx_rusty \| CFS \| Delta \| ---------o-----------o-----------o----------o Mean \| 7654.25% \| 7551.67% \| 1.11% \| ---------o-----------o-----------o----------o Variance \| 165.35714 \| 158.3333 \| 4.436% \| ---------o-----------o-----------o----------o o-----------o-----------o----------o \| rusty NUMA\| rusty ORIG\| Delta \| ---------o-----------o-----------o----------o Mean \| 7654.25% \| 7641.57% \| 0.1659% \| ---------o-----------o-----------o----------o Variance \| 165.35714 \| 1230.619 \| -86.5% \| ---------o-----------o-----------o----------o As expected, CPU util is quite a bit higher with scx_rusty than it is with CFS. Further experiments that could be interesting are always enabling direct-greedy stealing between domains within a NUMA node, and then comparing rusty NUMA and rusty ORIG. rusty NUMA prevents stealing between NUMA nodes, so this would show whether the locality introduced by NUMA awareness appropriately offsets the loss of work conservation. Major PFs --------- o-----------o-----------o----------o \| scx_rusty \| CFS \| Delta \| ---------o-----------o-----------o----------o Mean \| 5332 \| 3950 \| 36.566% \| ---------o-----------o-----------o----------o Variance \| 6975.5 \| 5986.333 \| 16.5237% \| ---------o-----------o-----------o----------o o-----------o-----------o----------o \| rusty NUMA\| rusty ORIG\| Delta \| ---------o-----------o-----------o----------o Mean \| 5332 \| 5336.5 \| -.084% \| ---------o-----------o-----------o----------o Variance \| 6975.5 \| 955.5 \| 630.03% \| ---------o-----------o-----------o----------o Also as expected, major page faults are far highe higher with scx_rusty than with CFS. This is expected even with NUMA awareness, given that scx_rusty is still less sticky than CFS. Further experiments that could be interesting are tuning the threshold for which we perform x NUMA migrations to try and keep this value even lower. The rate of major page faults between rusty NUMA and rusty ORIG were very close, though rusty NUMA was a bit lower. Signed-off-by: David Vernet <void@manifault.com>	2024-03-08 15:11:17 -06:00
David Vernet	0b1c3713b2	rusty: Remove lb_apply_weight param from lb_step() Let's just query self.tuner.fully_utilized directly and save a few lines of code. Signed-off-by: David Vernet <void@manifault.com>	2024-03-08 15:11:17 -06:00
David Vernet	758f762058	rusty: Move LoadBalancer out of rusty.rs More cleanup of scx_rusty. Let's move the LoadBalancer out of rusty.rs and into its own file. It will soon be extended quite a bit to support multi-NUMA and other multivariate LB cost functions, so it's time to clean things up and split it out. Signed-off-by: David Vernet <void@manifault.com>	2024-03-08 15:11:17 -06:00
David Vernet	94f75bcec6	rusty: Refactor Tuner and DomainGroup out of rusty.rs rusty.rs is growing a bit unwieldy. We're going to want to update its load balancing logic somewhat significantly to account for multi-NUMA and other cost functions, so let's start cleaning the code up so that things are more logically segmented and easier to work with. To start, we move the Tuner and DomainGroup/Domain objects into their own modules. Signed-off-by: David Vernet <void@manifault.com>	2024-03-08 15:10:37 -06:00
Jordan Rome	6c7617a037	Merge pull request #177 from jordalgo/libbpf-shallow Remove libbpf as a submodule	2024-03-08 06:12:40 -05:00
Jordan Rome	1769dece7d	Remove libbpf as a submodule Instead clone the libbpf repo at a specific hash during setup. This is to fix an issue whereby submodules are not included in the tarball and therefore won't be updated/fetched during setup after unzipping the tarball.	2024-03-07 18:31:09 -08:00
David Vernet	1a6ff1a871	Merge pull request #175 from sched-ext/docs [trivial] docs: Update rhone link	2024-03-07 10:00:57 -06:00
David Vernet	fdf5f5be55	docs: Update rhone link I changed by GitHub username to Byte-Lab, so let's update the docs. Signed-off-by: David Vernet <void@manifault.com>	2024-03-07 09:16:27 -06:00
Jordan Rome	a77793bd10	Merge pull request #174 from jordalgo/build-libbpf-static-only Libbpf - add BUILD_STATIC_ONLY flag	2024-03-05 19:31:26 -05:00
Jordan Rome	96fe285588	Libbpf - add BUILD_STATIC_ONLY flag	2024-03-05 15:11:51 -08:00
Jordan Rome	3eb700156a	Merge pull request #172 from jordalgo/libbpf-flags Always build libbpf as a PIE	2024-03-05 16:29:32 -05:00
Jordan Rome	38dab12459	Always build libbpf as a PIE This is to fix an error sometimes seen when compiling with gcc, whereby the position independent sched ext rust libraries don't play nicely with the libbpf library if it's not also built as position independent. This also explicitly sets the rustc relocation-mode to "pic", which is the default (just so this doesn't accidentally change out from under us).	2024-03-05 12:56:09 -08:00
Andrea Righi	a42dd32ff4	Merge pull request #173 from sched-ext/scx-rlfifo-warning scx_rlfifo: warn user about performance	2024-03-05 19:59:12 +01:00
Andrea Righi	be5e51dfaa	scx_rlfifo: print a performance warning banner scx_rlfifo is provided as a simple example to show how to use scx_rustland_core and it's not supposed to be used in a real production environment. To prevent performance bug reports print an explicit warning when it's started to clarify the goal of this scheduler. Signed-off-by: Andrea Righi <andrea.righi@canonical.com>	2024-03-05 19:36:17 +01:00
Andrea Righi	fe19754132	scx_rlfifo: replace 1ms sleep with sched_yield() Small improvement to make the scheduler a bit more responsive, without introducing too much complexity or too much CPU overhead. This can be achieved by replacing a sleep of 1ms with a sched_yield() every time that the scheduler has finished to dispatch all the queued tasks. This also makes the code a bit smaller and easier to read. Signed-off-by: Andrea Righi <andrea.righi@canonical.com>	2024-03-05 18:42:24 +01:00
Tejun Heo	db17905930	Merge pull request #170 from sched-ext/htejun meson-scripts/build_libbpf: Accommodate meson setting CC to "ccache $COMPILER"	2024-03-04 10:09:19 -10:00
Tejun Heo	069c390ef2	meson-scripts/build_libbpf: Accommodate meson setting CC to "ccache $COMPILER" Otherwise, we end up passing CC=ccache to libbpf's Makefile which triggers an error as ccache invoked on its own can't act as a stand-in for the compiler.	2024-03-04 10:04:25 -10:00

... 2 3 4 5 6 ...

639 Commits