description/open_source: update phoronix vs LWN to date 20230130

This commit is contained in:
Cheng Jian
2023-01-30 08:57:02 +08:00
parent 6610098823
commit ab743787e4
10 changed files with 194 additions and 7 deletions
+3
View File
@@ -145,6 +145,9 @@ $rq_nr_load = 1000 \times ux_nr + 100 \times top_nr + 10 \times fg_nr + bg_nr$
ColorOS 提供了 [MF(Multi Freearea)](https://github.com/oppo-source/android_kernel_modules_and_devicetree_oppo_sm8250/tree/oppo/sm8250_s_12.1/vendor/oplus/kernel/oplus_performance/multi_freearea) 提供了物理内存反碎片化的能力.
| 时间 | 作者 | 特性 | 描述 | 是否合入主线 | 链接 |
|:---:|:----:|:---:|:----:|:---------:|:----:|
| 2021/04/14 | lipeifeng@oppo.com <lipeifeng@oppo.com> | [mm: support multi_freearea to the reduction of external fragmentation](https://lore.kernel.org/all/20210414023803.937-1-lipeifeng@oppo.com) | TODO | v1 ☐☑✓ | [LORE](https://lore.kernel.org/all/20210414023803.937-1-lipeifeng@oppo.com) |
### 2.1.2 CSVM(Centralize Small Virtual Mem) 虚拟内存反碎片化机制
+2
View File
@@ -128,6 +128,7 @@ v5.7 引入了拆分锁检测的支持, 这依赖于 x86_64 intel CPU 遇到拆
### 1.2.2 Flexible Return and Event Delivery (FRED)
-------
[Linux 6.3 To Support Making Use Of Intel's New LKGS Instruction (Part Of FRED)](https://www.phoronix.com/news/Intel-LKGS-Linux-6.3)
| 时间 | 作者 | 特性 | 描述 | 是否合入主线 | 链接 |
|:---:|:----:|:---:|:----:|:---------:|:----:|
@@ -1069,6 +1070,7 @@ openEuler 提供了 [openEuler/prefetch_tuning](https://gitee.com/openeuler/pref
| 2021/12/24 | Huang Rui <ray.huang@amd.com> | [cpufreq: Introduce a new AMD CPU frequency control mechanism](https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/log/?id=38fec059bb69793f38cfa7a671d4bdbfe2a647aa) | TODO |v7 ☐☑✓ | [LORE v7,0/14](https://lore.kernel.org/all/20211224010508.110159-1-ray.huang@amd.com) |
| 2022/11/11 | Perry Yuan <Perry.Yuan@amd.com> | [Implement AMD Pstate EPP Driver](tps://lore.kernel.org/lkml/20221219064042.661122-1-perry.yuan@amd.com) | [AMD P-State EPP Patches Spun An 8th Time For Helping Out Linux Performance & Efficiency](https://www.phoronix.com/news/AMD-P-State-EPP-v8) | v4 ☐☑✓ 5.17-rc1 | [LORE v4,0/9](https://lore.kernel.org/all/20221110175847.3098728-1-Perry.Yuan@amd.com)<br>*-*-*-*-*-*-*-* <br>[LORE v8,00/13](https://lore.kernel.org/lkml/20221219064042.661122-1-perry.yuan@amd.com) |
| 2022/03/25 | Mario Limonciello <mario.limonciello@amd.com> | [Improve usability for amd-pstate](https://lore.kernel.org/all/20220325054228.5247-1-mario.limonciello@amd.com) | TODO | v1 ☐☑✓ | [LORE v1,0/3](https://lore.kernel.org/all/20220325054228.5247-1-mario.limonciello@amd.com)<br>*-*-*-*-*-*-*-* <br>[LORE v3,0/6](https://lore.kernel.org/linux-pm/20220414164801.1051-1-mario.limonciello@amd.com) |
| 2023/01/13 | Wyes Karny <wyes.karny@amd.com> | [amd_pstate: Add guided autonomous mode support](https://lore.kernel.org/all/20230113052141.2874296-1-wyes.karny@amd.com) | [AMD Updates P-State "Guided Autonomous Mode" Support For Linux](https://www.phoronix.com/news/AMD-Guided-Auto-Mode-v2) | v2 ☐☑✓ | [LORE v2,0/6](https://lore.kernel.org/all/20230113052141.2874296-1-wyes.karny@amd.com) |
+5
View File
@@ -243,8 +243,11 @@ $reclaim = current\_mem \times reclaim\_ratio \times max(0,1 \frac{psi_some}
## 8.2 getrandom vDSO
-------
[Implementing virtual system calls](https://lwn.net/Articles/615809)
[Linux Proposal Adding getrandom() To The vDSO For Better Performance](https://www.phoronix.com/news/Linux-getrandom-vDSO)
[A vDSO implementation of getrandom()](https://lwn.net/Articles/919008)
| 时间 | 作者 | 特性 | 描述 | 是否合入主线 | 链接 |
|:---:|:----:|:---:|:----:|:---------:|:----:|
@@ -898,6 +901,8 @@ Fedora 尝试优化 systemd 开机以及重启的时间, 参见 phoronix 报道
# 21 RUST 支持
-------
[Arm Helping With AArch64 Rust Linux Kernel Enablement](https://www.phoronix.com/news/AArch64-Rust-Linux-Kernel)
| 时间 | 作者 | 特性 | 描述 | 是否合入主线 | 链接 |
|:---:|:----:|:---:|:----:|:---------:|:----:|
| 2022/09/27 | Miguel Ojeda <ojeda@kernel.org> | [Rust support](https://lore.kernel.org/all/20220927131518.30000-1-ojeda@kernel.org) | TODO| v10 ☐☑✓ | [LORE 00/13](https://lore.kernel.org/all/20210414184604.23473-1-ojeda@kernel.org)<br>*-*-*-*-*-*-*-* <br>[LORE v10,0/27](https://lore.kernel.org/all/20220927131518.30000-1-ojeda@kernel.org) |
@@ -266,7 +266,7 @@ Linux 一开始是在一台i386上的机器开发的, i386 的硬件页表是2
| 时间 | 作者 | 特性 | 描述 | 是否合入主线 | 链接 |
|:----:|:----:|:---:|:----:|:---------:|:----:|
| 2022/09/21 | Huang Ying <ying.huang@intel.com> | [migrate_pages(): batch TLB flushing](https://lore.kernel.org/all/20220921060616.73086-1-ying.huang@intel.com) | 当前, migrate_pages()逐个迁移页面, 对每一页进行解除映射, 刷新 TLB, 然后恢复映射. 如果将多个页面传递给 migrate_pages(), 则有机会批量刷新和复制 TLB. TLB 冲洗 IPI 的总数可以大大减少. 并可以使用一些硬件加速器, 如 DSA 来加速页面复制. 因此, 在这个补丁中, 我们重构了 migrate_pages()实现, 并实现了 TLB 刷新批处理. 在此基础上, 可以实现硬件加速页面复制. 参见 phoronix 报道 [Intel Prepares Linux Batch TLB Flushing For Page Migration As A Big Performance Win](https://www.phoronix.com/news/Linux-Migrate-Pages-Batch-Flush). | v1 ☐☑✓ | [LORE v1,0/6](https://lore.kernel.org/all/20220921060616.73086-1-ying.huang@intel.com)<br>*-*-*-*-*-*-*-* <br>[LORE v1,0/8](https://lore.kernel.org/r/20221227002859.27740-1-ying.huang@intel.com) |
| 2022/09/21 | Huang Ying <ying.huang@intel.com> | [migrate_pages(): batch TLB flushing](https://lore.kernel.org/all/20220921060616.73086-1-ying.huang@intel.com) | 当前, migrate_pages()逐个迁移页面, 对每一页进行解除映射, 刷新 TLB, 然后恢复映射. 如果将多个页面传递给 migrate_pages(), 则有机会批量刷新和复制 TLB. TLB 冲洗 IPI 的总数可以大大减少. 并可以使用一些硬件加速器, 如 DSA 来加速页面复制. 因此, 在这个补丁中, 我们重构了 migrate_pages()实现, 并实现了 TLB 刷新批处理. 在此基础上, 可以实现硬件加速页面复制. 参见 phoronix 报道 [Intel Prepares Linux Batch TLB Flushing For Page Migration As A Big Performance Win](https://www.phoronix.com/news/Linux-Migrate-Pages-Batch-Flush) 和 [Intel Optimization Around Batched TLB Flushing For Folios Looks Great](https://www.phoronix.com/news/Migrate-Pages-Batch-TLB-Flush-F). | v1 ☐☑✓ | [LORE v1,0/6](https://lore.kernel.org/all/20220921060616.73086-1-ying.huang@intel.com)<br>*-*-*-*-*-*-*-* <br>[LORE v1,0/8](https://lore.kernel.org/r/20221227002859.27740-1-ying.huang@intel.com)<br>*-*-*-*-*-*-*-* <br>[LORE v3,0/9](https://lore.kernel.org/all/20230116063057.653862-1-ying.huang@intel.com) |
| 2022/10/28 | Yicong Yang <yangyicong@huawei.com> | [arm64: support batched/deferred tlb shootdown during page reclamation](https://patchwork.kernel.org/project/linux-mm/cover/20221028081255.19157-1-yangyicong@huawei.com/) | 虽然 ARM64 有硬件进行 tlb shootdown, 但是使用 tlbi 进行硬件广播的开销也不小. 一个最简单的微基准测试显示, 即使在只有 8 核的骁龙 888 上, 即使只分页一个进程映射的页面, ptep_clear_flush() 的开销(perf top) 也达到了 5.36%. 当页面由多个进程映射或 HW 有更多 CPU 时, 由于 tlb shootdown 糟糕的可伸缩性, 成本应该会变得更高, 同样的基准测试在 100 核左右的 ARM64 服务器上可以导致 16.99% 的 CPU 消耗. 这个补丁集利用现有的 BATCHED_UNMAP_TLB_FLUSH. 只在第一阶段 arch_tlbbatch_add_mm() 中发送 tlbi 指令. 在骁龙 888 上的测试表明, 补丁集消除了 ptep_clear_flush() 的开销. 在骁龙 888 上, 即使是单个进程映射的一个页面, 微基准测试也要快 5%. 有了这个支持, 我们可以对内存回收和[迁移](https://lore.kernel.org/lkml/20220921060616.73086-1-ying.huang@intel.com)做更多的优化. | v5 ☐☑ | [LORE 0/4](https://lore.kernel.org/lkml/20220707125242.425242-1-21cnbao@gmail.com)<br>*-*-*-*-*-*-*-* <br>[LORE v5,0/2](https://lore.kernel.org/r/20221028081255.19157-1-yangyicong@huawei.com)<br>*-*-*-*-*-*-*-* <br>[LORE v6,0/2](https://lore.kernel.org/r/20221115031425.44640-1-yangyicong@huawei.com)<br>*-*-*-*-*-*-*-* <br>[LORE v7,0/2](https://lore.kernel.org/r/20221117082648.47526-1-yangyicong@huawei.com) |
@@ -5430,7 +5430,7 @@ sys_fork
#### 8.2.2.3 页表的写时拷贝
-------
[Introduce Copy-On-Write to Page Table](https://patchwork.kernel.org/project/linux-mm/cover/20220927162957.270460-1-shiyn.lin@gmail.com) 这组补丁集为 PTE 级页表引入了写入时拷贝(COW).
[Introduce Copy-On-Write to Page Table](https://patchwork.kernel.org/project/linux-mm/cover/20220927162957.270460-1-shiyn.lin@gmail.com) 这组补丁集为 PTE 级页表引入了写入时拷贝(COW). 参见 [Memory-management short topics: page-table sharing and working sets](https://lwn.net/Articles/919143)
在用户需要程序副本才能在隔离环境中运行的情况下, COW PTE 提高了性能. 基于反馈的模糊器 (例如, AFL) 和微服务框架是两个主要的例子. 例如, COW PTE 在 fuzzer(AFL) 上运行 SQLite 时, 吞吐量增加了 9.3 倍.
@@ -102,6 +102,7 @@
| 6.0 | NA | NA | [Linux 6.0 Supporting New Intel/AMD Hardware, Performance Improvements & Much More](https://www.phoronix.com/review/linux-60-features), [6.0-rc1](https://www.phoronix.com/news/Linux-6.0-rc1-Released) |
| 6.1 | NA | NA | [Linux 6.1 Features Include Initial Rust Code, MGLRU, New AMD CPU Features, More Security](https://www.phoronix.com/review/linux-61-features), [The Most Interesting New Features For Linux 6.1](https://www.phoronix.com/news/Linux-6.1-Features) |
| 6.2 | NA | NA | [The Many New Features On The Horizon For Linux 6.2](https://www.phoronix.com/news/Linux-6.2-Early-Features)<br>*-*-*-*-*-*-*-* <br>[Linux 6.2-rc1 Brings Stable Intel Arc Graphics, Call Depth Tracking & Many More Features](https://www.phoronix.com/news/Linux-6.2-rc1-Released)<br>*-*-*-*-*-*-*-* <br>[Linux 6.2 Features: Stable Intel Arc Graphics. RTX 30 Support, Intel On Demand + IFS Ready](https://www.phoronix.com/review/linux-62-features) |
| 6.3 | NA | NA | [Linux 6.3 Features Expected From AMD Auto IBRS To Pluton CRB TPM2 & Dropping Old Code](https://www.phoronix.com/news/Linux-6.3-Early-Features-Look) |
# 6 业界会议
-------
+19 -4
View File
@@ -331,7 +331,7 @@ SCHED_IDLE 跟 SCHED_BATCH 一样, 是 CFS 中的一个策略, SCHED\_IDLE 的
| 2022/02/17 | Abel Wu <wuyun.abel@bytedance.com> | [introduce sched-idle balancing](https://lore.kernel.org/all/20220217154403.6497-1-wuyun.abel@bytedance.com) | 当前负载平衡主要基于 cpu capacity 和 task util, 这在整体吞吐量的 POV 中是有意义的. 虽然如果存在 sched 闲置或闲置 RQ, 则可以通过减少过载 CFS RQ 的数量来完成一些改进. 当 CFS RQ 上有多个可伸缩的非闲置任务时(因为 schedidle CPU 被视为闲置 CPU), CFS RQ 被认为是过载的. 空闲任务计入 rq->cfs.idle_h_nr_running.<br> 过载的 CFS RQ 可能会导致两种任务类型的性能问题:<br>1. 对于诸如 SCHED_NORMAL 之类的延迟关键任务, RQ 中的等待时间将增加并导致更高的 PCT99 延迟, 并且如果存在 SCHED_DILE, 批处理任务 SCHED_BATCH 可能无法充分利用 CPU 容量, 因此吞吐量较差.<br> 所以简而言之, sched-idle balancing 的目标是让非闲置任务充分利用 CPU 资源.<br> 为此, 我们主要做两件事:<br>1. 为 sched-idle 的 CPU 拉取 non-idle 的任务来运行, 或者将 overload CPU 上的任务拉取到 idle 的 CPU 上.<br>2. 防止在 RQ 中 PULL 出最后一个非闲置任务. 此外 overloaded CPUs 的掩码会周期性更新, 空闲路径在 LLC 域上. 这个 cpumask 还将在 SIS 中用作过滤器, 改善空闲的 CPU 搜索. | v1 ☐☑✓ | [LORE v1,0/5](https://lore.kernel.org/all/20220217154403.6497-1-wuyun.abel@bytedance.com)<br>*-*-*-*-*-*-*-* <br>[LORE v2,0/2](https://lore.kernel.org/lkml/20220409135104.3733193-1-wuyun.abel@bytedance.com) |
| 2022/08/09 | zhangsong <zhangsong34@huawei.com> | [sched/fair: Introduce priority load balance for CFS]](https://lore.kernel.org/all/20220809132945.3710583-1-zhangsong34@huawei.com) | 对于 NORMAL 和 IDLE 任务的共存, 当 CFS 触发负载均衡时, 将 NORMAL(Latency Sensitive) 任务从繁忙的 src CPU 迁移到 dst CPU, 最后迁移 IDLE 任务是合理的. 这对于减少 SCHED_IDLE 任务的干扰非常重要.<br> 但是当前的 cfs_tasks 链表同时包含了 NORMAL 任务和 SCHED_IDLE 等任务, 且没有按照优先级进行排序, 因此无法保证能及时从 busiest 的等待队列中拉出一定数量的正常任务而不是空闲任务 <br> 因此需要将 cfs_tasks 分成两个不同的列表, 并确保非空闲列表中的任务能够首先迁移. 该补丁引入 cfs_idle_tasks 链表维护 SCHED_IDLE 的任务, 原来的 cfs_tasks 只维护 SCHED_NORMAL 的任务. 负载均衡时优先迁移 SCHED_NORMAL 的任务.<br> 测试发现: 少量的 NORMAL 任务与大量的 IDLE 任务搭配, 通过该补丁, NORMAL 任务延迟较当前降低约 5~10%. | v1 ☐☑✓ | [LORE](https://lore.kernel.org/all/20220809132945.3710583-1-zhangsong34@huawei.com)<br>*-*-*-*-*-*-*-* <br>[LORE v2](https://lore.kernel.org/lkml/20220810015636.3865248-1-zhangsong34@huawei.com)<br>*-*-*-*-*-*-*-* <br>[LORE v3](https://lore.kernel.org/lkml/20220810092546.3901325-1-zhangsong34@huawei.com))<br>*-*-*-*-*-*-*-* <br>[LORE v4](https://lore.kernel.org/all/20221102035301.512892-1-zhangsong34@huawei.com) |
| 2022/08/25 | Vincent Guittot <vincent.guittot@linaro.org> | [sched/fair: fixes in presence of lot of sched_idle tasks](https://lore.kernel.org/all/20220825122726.20819-1-vincent.guittot@linaro.org) | TODO | v1 ☐☑✓ | [LORE v1,0/4](https://lore.kernel.org/all/20220825122726.20819-1-vincent.guittot@linaro.org) |
| 2022/10/03 | Vincent Guittot <vincent.guittot@linaro.org> | [sched/fair: limit sched slice duration](https://lore.kernel.org/all/20221003122111.611-1-vincent.guittot@linaro.org) | TODO | v3 ☐☑✓ | [LORE](https://lore.kernel.org/all/20221003122111.611-1-vincent.guittot@linaro.org) |
| 2022/10/03 | Vincent Guittot <vincent.guittot@linaro.org> | [sched/fair: limit sched slice duration](https://lore.kernel.org/all/20221003122111.611-1-vincent.guittot@linaro.org) | TODO | v3 ☐☑✓ | [LORE](https://lore.kernel.org/all/20221003122111.611-1-vincent.guittot@linaro.org)<br>*-*-*-*-*-*-*-* <br>[LORE v4](https://lore.kernel.org/all/20230113133613.257342-1-vincent.guittot@linaro.org) |
#### 1.1.5.4 cgroup SCHED_IDLE support
@@ -548,6 +548,7 @@ linux 调度器定义了多个调度类, 不同调度类的调度优先级不同
| [超线程的两个线程资源是动态分配的还是固定一半一半的?](https://www.zhihu.com/question/59721493) | NA |
| [英特尔超线程技术](https://baike.baidu.com/item/英特尔超线程技术/10233952) | NA |
| [为什么cinebench r15和r20在CPU满载渲染时超线程可以显著提高跑分?](https://www.zhihu.com/question/319200765/answer/646240231) | 介绍了 TOPDOWN 以及 Intel Vtune 工具 |
| [Will Hyper-Threading Improve Processing Performance?](https://www.dasher.com/will-hyper-threading-improve-processing-performance) | 解释 SMT 如何提升系统的性能 |
#### 1.5.4.1 SMT aware
-------
@@ -1352,6 +1353,7 @@ X86 下提供了一种 Fake Numa 的方式来模拟 NUMA 配置.
|:---:|:----:|:---:|:----:|:---------:|:----:|
| 2022/07/19 | Tariq Toukan <tariqt@nvidia.com> | [Introduce and use NUMA distance metrics](https://lore.kernel.org/all/20220719162339.23865-1-tariqt@nvidia.com) | NVIDIA 工程师一直在 Linux 内核中研究 NUMA 距离指标, 以取代某些驱动程序目前用于 NUMA 感知内存分配的简单本地 / 远程 NUMA 首选项接口. 在他们的测试中, 这种改进的 NUMA 距离处理对吞吐量和 CPU 利用率产生了 "显著的性能影响". 根据调度程序的 sched_numa_find_closest() 实现并公开 CPU spread API sched_cpus_set_spread(). 在给定 NUMA 节点的情况下, 基于距离设置 CPU 分布替代基于 cpumask_local_spread() 的传统逻辑. 在 mlx5 和 enic 设备驱动程序中使用它. 这将使得 NUMA 首选项 (本地 / 远程) 替换为考虑实际距离的改进首选项, 因此短距离的远程 NUMA 优先于较远的 NUMA. | v3 ☐☑✓ | [LORE v3,0/3](https://lore.kernel.org/all/20220719162339.23865-1-tariqt@nvidia.com)<br>*-*-*-*-*-*-*-* <br>[2022/07/28 LORE v4,0/3](20220728191203.4055-1-tariqt@nvidia.com) |
| 2022/08/17 | Valentin Schneider <vschneid@redhat.com> | [cpumask, sched/topology: NUMA-aware CPU spreading interface](https://lore.kernel.org/all/20220816180727.387807-1-vschneid@redhat.com) | TODO | v1 ☐☑✓ | [2022/08/16 LORE v1,0/5](https://lore.kernel.org/all/20220816180727.387807-1-vschneid@redhat.com)<br>*-*-*-*-*-*-*-* <br>[2022/08/17 LORE v2,0/5](https://lore.kernel.org/lkml/20220817175812.671843-1-vschneid@redhat.com)<br>*-*-*-*-*-*-*-* <br>[LORE v5,0/3](https://lore.kernel.org/all/20221021121927.2893692-1-vschneid@redhat.com) |
| 2023/01/20 | Yury Norov <yury.norov@gmail.com> | [sched: cpumask: improve on cpumask_local_spread() locality](https://lore.kernel.org/all/20230121042436.2661843-1-yury.norov@gmail.com) | cpumask_local_spread() 当前检查本地节点是否存在第 i 个 CPU, 如果没有发现, 则在所有非本地 CPU 之间进行平面搜索. 我们可以通过检查每个 NUMA 跳的 CPU 来做得更好. 这对 NUMA 机器有显著的性能影响, 例如, 当使用 NUMA 感知的分配内存和 NUMA 感知 IRQ 关联提示时. | v1 ☐☑✓ | [LORE v1,0/9](https://lore.kernel.org/all/20230121042436.2661843-1-yury.norov@gmail.com) |
## 4.2 负载均衡总概
@@ -2566,9 +2568,14 @@ commit [6e5fb223e89d ("mm: sched: numa: Implement constant, per task Working Set
#### 4.6.3.4 Process Adaptive
-------
现有的扫描周期机制涉及从每线程统计数据导出的扫描周期, 这样的开销很大.
在之前讨论过程中, Mel 提出了几个增强当前 NUMA Balancing 的想法. 其中一个建议如下: 跟踪哪些线程访问 VMA. 建议使用 unsigned long pid_mask, 并使用较低的位来标记大约哪些线程访问 VMA. 跳过未捕获故障的 vma. 由于 PID 冲突, 这将是近似的, 但会减少线程不感兴趣的区域的扫描. 上述建议不会惩罚对 vma 不感兴趣的线程, 从而减少扫描开销.
| 时间 | 作者 | 特性 | 描述 | 是否合入主线 | 链接 |
|:----:|:----:|:---:|:---:|:----------:|:----:|
| 2022/01/28 | Bharata B Rao <bharata@amd.com> | [sched/numa: Process Adaptive autoNUMA](https://lore.kernel.org/lkml/20220128052851.17162-1-bharata@amd.com) | 实现了一种进程自适应 autoNUMA 算法 (Process Adaptive autoNUMA, PAN), 用于计算 autoNUMA 扫描周期.<br> 在现有的扫描周期计算机制中: 1. 扫描周期是从每线程的统计数据中派生出来的. 2. 静态阈值(NUMA_PERIOD_threshold) 用于更改扫描速率.<br> 这组补丁集将 NUMA fault 按照不同的维护划分, 如本地的与远程的 (local vs. remote), 私有的和共享的(private vs. shared). 然后在每个进程级别收集 numa faults 统计数据, 从而更好地捕获应用程序行为. 不再使用静态阈值, 而是根据远程故障率来学习和调整扫描速率, 可以更好地响应不同的工作负载行为. 由于进程的线程已经被视为一个 numa_group, 因此我们在任务的[内存管理] 中添加了一组度量标准, 以跟踪各种类型的错误并从中推导出扫描速度. 新的每进程故障统计数据只对每进程扫描周期计算有贡献, 而现有的每线程统计数据继续对 numa_group 统计数据有贡献, 后者最终确定跨节点迁移内存和线程的阈值. 参见 phoronix 的报道 [AMD Cooking Up A"PAN"Feature That Can Help Boost Linux Performance](https://www.phoronix.com/scan.php?page=news_item&px=AMD-PAN-Linux-RFC) | v0 ☐ | [LKML v0,0/5](https://lkml.org/lkml/2022/1/28/16), [LORE](https://lore.kernel.org/lkml/20220128052851.17162-1-bharata@amd.com) |
| 2022/01/28 | Bharata B Rao <bharata@amd.com> | [sched/numa: Process Adaptive autoNUMA](https://lore.kernel.org/lkml/20220128052851.17162-1-bharata@amd.com) | 实现了一种进程自适应 autoNUMA 算法 (Process Adaptive autoNUMA, PAN). 在每个进程级别上收集 NUMA 故障统计信息, 以更好地捕获应用程序行为, 计算 autoNUMA 扫描周期.<br> 在现有的扫描周期计算机制中: 1. 扫描周期是从每线程的统计数据中派生出来的. 2. 静态阈值(NUMA_PERIOD_threshold) 用于更改扫描速率.<br> 这组补丁集将 NUMA fault 按照不同的维护划分, 如本地的与远程的 (local vs. remote), 私有的和共享的(private vs. shared). 然后在每个进程级别收集 numa faults 统计数据, 从而更好地捕获应用程序行为. 不再使用静态阈值, 而是根据远程故障率来学习和调整扫描速率, 可以更好地响应不同的工作负载行为. 由于进程的线程已经被视为一个 numa_group, 因此我们在任务的[内存管理] 中添加了一组度量标准, 以跟踪各种类型的错误并从中推导出扫描速度. 新的每进程故障统计数据只对每进程扫描周期计算有贡献, 而现有的每线程统计数据继续对 numa_group 统计数据有贡献, 后者最终确定跨节点迁移内存和线程的阈值. 参见 phoronix 的报道 [AMD Cooking Up A"PAN"Feature That Can Help Boost Linux Performance](https://www.phoronix.com/scan.php?page=news_item&px=AMD-PAN-Linux-RFC) | v0 ☐ | [LKML v0,0/5](https://lkml.org/lkml/2022/1/28/16), [LORE](https://lore.kernel.org/lkml/20220128052851.17162-1-bharata@amd.com) |
| 2023/01/16 | Raghavendra K T <raghavendra.kt@amd.com> | [sched/numa: Enhance vma scanning](https://lore.kernel.org/all/cover.1673610485.git.raghavendra.kt@amd.com) | 借助了 Mel 的建议核想法, 不同于 Process Adaptive autoNUMA. 本补丁集<br>1. 最多跟踪 4 个最近访问 vma 的线程, 只扫描访问 vma 的线程. (注意: 只使用 unsigned int. 实验表明, 追踪 8 种不同的 pid 开销更大)<br>2. 前 2 次无条件允许线程扫描 vmas, 以保持扫描的初衷.<br>3. 如果有超过 4 个线程(即超过我们可以记住的 pid), 默认允许扫描, 因为我们可能会错过记录当前线程是否对 vma 有任何兴趣.<br>通过这个补丁集, 可以看到扫描开销(AutoNuma 开销) 大幅减少, 其中一些 enchmark 提高了性能, 而其他的几乎没有倒退. | v1 ☐☑✓ | [LORE v1,0/1](https://lore.kernel.org/all/cover.1673610485.git.raghavendra.kt@amd.com) |
### 4.6.4 NUMA Balancing Placement And Migration
@@ -5782,6 +5789,7 @@ Intel 的 [Wult/Wake Up Latency Tracer](https://github.com/intel/wult) 一个在
| 时间 | 作者 | 特性 | 描述 | 是否合入主线 | 链接 |
|:----:|:----:|:---:|:----:|:---------:|:----:|
| 2022/10/31 | Zhang Qiao <zhangqiao22@huawei.com> | [sched: sched_fork() optimizations](https://lore.kernel.org/all/20221031125113.72980-1-zhangqiao22@huawei.com) | sched_fork() 使用当前 CPU 初始化新任务的 vruntime, 但新任务可能不在这个 CPU 上运行. 所以这个补丁集将解决这个问题. | v1 ☐☑✓ | [LORE v1,0/2](https://lore.kernel.org/all/20221031125113.72980-1-zhangqiao22@huawei.com)<br>*-*-*-*-*-*-*-* <br>[LORE v2,0/2](https://lore.kernel.org/all/20221103120720.39873-1-zhangqiao22@huawei.com) |
| 2023/01/27 | Roman Kagan <rkagan@amazon.de> | [sched/fair: sanitize vruntime of entity being placed](https://lore.kernel.org/all/20230127163230.3339408-1-rkagan@amazon.de) | 当一个调度实体被放置到 cfs_rq 上时, 它的 vruntime 被拉到 cfs_rq->min_vruntime 附近, 这样实体在向后放置时不会获得额外的提升. 然而, 如果被放置的实体很长时间没有执行, 它的 vruntime 可能会落后太多 (例如, 当 cfs_rq 执行一个低权重的 hog 时), 这可能会由于 s64 溢出而导致 vruntime 比较相反. 这将导致实体以其原始 vruntime 的方式向前放置, 因此它将永远不会有效地被运行. 为了防止这种情况, 如果实体的执行时间没有比特征计算器时间尺度长得多, 请忽略它的运行时间. | v1 ☐☑✓ | [LORE](https://lore.kernel.org/all/20230127163230.3339408-1-rkagan@amazon.de) |
### 10.1.2 shared page tables
@@ -5857,6 +5865,13 @@ c. 在多个进程之间共享(例如进程会话).
| 2022/09/08 | Pavel Tikhomirov <ptikhomirov@virtuozzo.com> | [Add CABA tree to task_struct](https://patchwork.kernel.org/project/linux-mm/patch/20220908220944.822942-1-ptikhomirov@virtuozzo.com) | 675404 | v4 ☐☑ | [LORE v4,0/2](https://lore.kernel.org/all/20220908220944.822942-1-ptikhomirov@virtuozzo.com) |
### 10.1.4 进程执行
-------
| 2023/01/19 | Giuseppe Scrivano <gscrivan@redhat.com> | [exec: add PR_HIDE_SELF_EXE prctl](https://lore.kernel.org/all/20230119170718.3129938-1-gscrivan@redhat.com) | 隐藏进程的 exe 程序防止容器内的安全问题. 类似于 CVE-2019-5736. 参见 LWN 报道 [Hiding a process's executable from itself](https://lwn.net/Articles/920384) | v2 ☐☑✓ | [LORE v2,0/2](https://lore.kernel.org/all/20230119170718.3129938-1-gscrivan@redhat.com) |
## 10.3 IPC
-------
@@ -5986,7 +6001,7 @@ Google 的 Peter Oskolkov 发布了[最早的 RFC v0.1 补丁](https://lore.kern
|:----:|:----:|:---:|:----:|:---------:|:----:|
| 2021/12/14 | Peter Oskolkov <posk@google.com>/<posk@posk.io> | [sched,mm,x86/uaccess: implement User Managed Concurrency Groups](https://lore.kernel.org/patchwork/cover/1433967) | UMCG (User-Managed Concurrency Groups) | [PatchWork RFC,v0.1,0/9](https://lore.kernel.org/patchwork/cover/1433967)<br>*-*-*-*-*-*-*-* <br>[2021/07/08 PatchWork RFC,0/3,v0.2](https://lore.kernel.org/patchwork/cover/1455166)<br>*-*-*-*-*-*-*-* <br>[2021/07/16 PatchWork RFC,0/4,v0.3](https://lore.kernel.org/patchwork/cover/1461708)<br>*-*-*-*-*-*-*-* <br>[2021/08/01 PatchWork 0/4,v0.4](https://lore.kernel.org/patchwork/cover/1470650)<br>*-*-*-*-*-*-*-* <br>[2021/08/01 LWN 0/4,v0.5](https://lore.kernel.org/patchwork/cover/1470650)<br>*-*-*-*-*-*-*-* <br>[2021/10/12 PatchWork v0.7,0/5](https://patchwork.kernel.org/project/linux-mm/cover/20211012232522.714898-1-posk@google.com)<br>*-*-*-*-*-*-*-* <br>[2021/11/04 PatchWork v0.8,0/6](https://patchwork.kernel.org/project/linux-mm/cover/20211104195804.83240-1-posk@google.com)<br>*-*-*-*-*-*-*-* <br>[2021/11/21 PatchWork v0.9,0/6](https://patchwork.kernel.org/project/linux-mm/cover/20211121212040.8649-1-posk@google.com)<br>*-*-*-*-*-*-*-* <br>[2021/11/23 PatchWork v0.9.1,0/6](https://patchwork.kernel.org/project/linux-mm/cover/20211122211327.5931-1-posk@google.com) |
| 2022/01/20 | Paul Gortmaker <paul.gortmaker@windriver.com> | [sched: User Managed Concurrency Groups](https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/log/?id=abedf8e2419fb873d919dd74de2e84b510259339) | Peter Zijlstra 对 UMCG 的重新实现. | v9 ☑ 4.6-rc1 | [PatchWork RFC,0/3](https://patchwork.kernel.org/project/linux-mm/cover/20211214204445.665580974@infradead.org)<br>*-*-*-*-*-*-*-* <br>[PatchWork RFC,v2,0/5](https://patchwork.kernel.org/project/linux-mm/cover/20220120155517.066795336@infradead.org) |
| 2022/10/19 | Andrei Vagin <avagin@gmail.com> | [seccomp: add the synchronous mode for seccomp_unotify](https://lore.kernel.org/all/20221020011048.156415-1-avagin@gmail.com) | seccomp_unotify 允许特权更大的进程代表特权更小的进程执行操作. 在许多情况下, 工作流是完全同步的. 它意味着一个目标进程触发一个系统调用, 并将控制传递给一个主管进程, 后者处理系统调用并将控制返回给目标进程. 在这个上下文中, "同步" 意味着只有一个进程在运行, 另一个正在等待.<br> 新的 WF_CURRENT_CPU 标志建议调度器将唤醒对象移动到当前 CPU. 对于这样的同步工作流, 它使上下文切换速度提高了几倍. 测试发现, 原来每个相互作用需要 12µs, 借助这个补丁这个过程只需要 3µs. | v2 ☐☑✓ | [LORE v2,0/5](https://lore.kernel.org/all/20221020011048.156415-1-avagin@gmail.com) |
| 2022/10/19 | Andrei Vagin <avagin@gmail.com> | [seccomp: add the synchronous mode for seccomp_unotify](https://lore.kernel.org/all/20221020011048.156415-1-avagin@gmail.com) | seccomp_unotify 允许特权更大的进程代表特权更小的进程执行操作. 在许多情况下, 工作流是完全同步的. 它意味着一个目标进程触发一个系统调用, 并将控制传递给一个主管进程, 后者处理系统调用并将控制返回给目标进程. 在这个上下文中, "同步" 意味着只有一个进程在运行, 另一个正在等待.<br> 新的 WF_CURRENT_CPU 标志建议调度器将唤醒对象移动到当前 CPU. 对于这样的同步工作流, 它使上下文切换速度提高了几倍. 测试发现, 原来每个相互作用需要 12µs, 借助这个补丁这个过程只需要 3µs. | v2 ☐☑✓ | [LORE v2,0/5](https://lore.kernel.org/all/20221020011048.156415-1-avagin@gmail.com)<br>*-*-*-*-*-*-*-* <br>[LORE v4,0/6](https://lore.kernel.org/lkml/20230124234156.211569-1-avagin@google.com) |
### 11.2.2 基于 eBPF 的可编程调度框架
@@ -6027,7 +6042,7 @@ Roman Gushchin 在邮件列表发起了 BPF 对调度器的潜在应用的讨论
| 时间 | 作者 | 特性 | 描述 | 是否合入主线 | 链接 |
|:----:|:----:|:---:|:----:|:---------:|:----:|
| 2021/09/15 | Roman Gushchin <guro@fb.com> | [Scheduler BPF](https://www.phoronix.com/scan.php?page=news_item&px=Linux-BPF-Scheduler) | NA | RFC ☐ | [PatchWork rfc,0/6](https://patchwork.kernel.org/project/netdevbpf/cover/20210916162451.709260-1-guro@fb.com)<br>*-*-*-*-*-*-*-* <br>[LPC 2021](https://linuxplumbersconf.org/event/11/contributions/954)<br>*-*-*-*-*-*-*-* <br>[LKML](https://lkml.org/lkml/2021/9/16/1049), [LWN](https://lwn.net/Articles/869433), [LWN](https://lwn.net/Articles/873244) |
| 2022/11/29 | Tejun Heo <tj@kernel.org> | [sched: Implement BPF extensible scheduler class](https://lore.kernel.org/all/20221130082313.3241517-1-tj@kernel.org) | 随后 FaceBook 进一步扩展, 引入 sched_ext 模块, 使用 eBPF 对调度器进行可编程重构. [Experimental Patches Allow eBPF To Extend The Linux Kernel's Scheduler](https://www.phoronix.com/news/RFC-eBPF-Linux-Scheduler), [The BPF extensible scheduler class](https://lwn.net/Articles/916291). | v1 ☐☑✓ | [LORE](https://lore.kernel.org/all/20221130082313.3241517-1-tj@kernel.org) |
| 2022/11/29 | Tejun Heo <tj@kernel.org> | [sched: Implement BPF extensible scheduler class](https://lore.kernel.org/all/20221130082313.3241517-1-tj@kernel.org) | 随后 FaceBook 进一步扩展, 引入 sched_ext 模块, 使用 eBPF 对调度器进行可编程重构. [Experimental Patches Allow eBPF To Extend The Linux Kernel's Scheduler](https://www.phoronix.com/news/RFC-eBPF-Linux-Scheduler), [The BPF extensible scheduler class](https://lwn.net/Articles/916291), [Patches Updated For Hooking eBPF Programs Into The Linux Kernel Scheduler](https://www.phoronix.com/news/Linux-Scheduler-eBPF-v2-sched). | v1 ☐☑✓ | [LORE](https://lore.kernel.org/all/20221130082313.3241517-1-tj@kernel.org)<br>*-*-*-*-*-*-*-* <br>[LORE v2,00/30](https://lore.kernel.org/lkml/20230128001639.3510083-1-tj@kernel.org) |
#### 11.2.2.2 Google 的 ghOSt
+121 -1
View File
@@ -204,4 +204,124 @@ https://www.latexlive.com
MGLRU 合入后, 引起了不少场景的性能劣化, 参见 phoronix 报道 [An MGLRU Performance Regression Fix Is On The Way Plus Another Optimization](https://www.phoronix.com/news/MGLRU-SVT-Performance-Fix).
[Linux 6.2 Features: Stable Intel Arc Graphics. RTX 30 Support, Intel On Demand + IFS Ready](https://www.phoronix.com/review/linux-62-features)
| 2020/02/27 | Valentin Schneider <valentin.schneider@arm.com> | [sched, arm64: enable CONFIG_SCHED_SMT for arm64](https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/log/?id=6f693dd5be08237b337f557c510d99addb9eb9ec) | TODO | v2 ☑✓ 5.7-rc1 | [LORE v2,0/2](https://lore.kernel.org/all/20200227191433.31994-1-valentin.schneider@arm.com) |
| 2022/12/02 | Brian Foster <bfoster@redhat.com> | [proc: improve root readdir latency with many threads](https://lore.kernel.org/all/20221202171620.509140-1-bfoster@redhat.com) | TODO | v3 ☐☑✓ | [LORE v3,0/5](https://lore.kernel.org/all/20221202171620.509140-1-bfoster@redhat.com) |
| 2023/01/09 | Yian Chen <yian.chen@intel.com> | [Enable LASS (Linear Address space Separation)](https://lore.kernel.org/all/20230110055204.3227669-1-yian.chen@intel.com) | 参见 LWN 报道 [Support for Intel's LASS](https://lwn.net/Articles/919683) 和 phoronix 报道 [Intel Posts Linux Patches For Linear Address Space Separation (LASS)](https://www.phoronix.com/news/Linear-Address-Space-Separation) | v1 ☐☑✓ | [LORE v1,0/7](https://lore.kernel.org/all/20230110055204.3227669-1-yian.chen@intel.com) |
[[LSF/MM/BFP TOPIC] Storage: Copy Offload](https://lkml.kernel.org/linux-block/f0e19ae4-b37a-e9a3-2be7-a5afb334a5c3@nvidia.com)
[LSFMM: Copy offload](https://lwn.net/Articles/548347)
[Storage: Xcopy Offload](https://blog.csdn.net/flyingnosky/article/details/123533554)
| 时间 | 作者 | 特性 | 描述 | 是否合入主线 | 链接 |
|:---:|:----:|:---:|:----:|:---------:|:----:|
| 2014/05/28 | Martin K. Petersen <martin.petersen@oracle.com> | [Copy offload](https://lore.kernel.org/all/1401335565-29865-1-git-send-email-martin.petersen@oracle.com) | TODO | v1 ☐☑✓ | [LORE](https://lore.kernel.org/all/1401335565-29865-1-git-send-email-martin.petersen@oracle.com) |
| 2022/11/23 | Nitesh Shetty <nj.shetty@samsung.com> | [Implement copy offload support](https://lore.kernel.org/all/20221123055827.26996-1-nj.shetty@samsung.com) | TODO | v5 ☐☑✓ | [LORE v5,0/10](https://lore.kernel.org/all/20221123055827.26996-1-nj.shetty@samsung.com) |
[调度器 34—RT 负载均衡](https://www.cnblogs.com/hellokitty2/p/15974333.html)
[实时调度负载均衡](https://github.com/freelancer-leon/notes/blob/master/kernel/sched/sched_rt_load_balance.md)
[Latencies, schedulers, interrupts oh my! The epic story of a Linux Kernel upgrade](https://www.nutanix.dev/2021/12/09/latencies-schedulers-interrupts-oh-my-the-epic-story-of-a-linux-kernel-upgrade)
[RISC-V Hibernation Support / Suspend-To-Disk Nears The Linux Kernel](https://www.phoronix.com/news/RISC-V-Hibernation-Linux)
[Intel Preparing New Linux"PerfMon"Performance Monitoring Support For IOMMU](https://www.phoronix.com/news/Intel-IOMMU-VT-d-4.0-PerfMon)
[ARM64 手动搭建 kdump 环境](https://blog.csdn.net/m0_37797953/article/details/107491356)
[crash 命令 —— list](https://www.cnblogs.com/pengdonglin137/p/16046328.html)
[CRASH 安装和调试](https://www.cnblogs.com/Linux-tech/p/14110330.html)
[fujitsu/crash-trace](https://github.com/fujitsu/crash-trace)
[How to display or retrieve ftrace data from the kernel crash dump?](https://access.redhat.com/solutions/239433)
5.7-rc1 [psi: Optimize switching tasks inside shared cgroups](https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=36b238d5717279163859fb6ba0f4360abcafab83)
5.13-rc1 [psi: Optimize task switch inside shared cgroups](https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=4117cebf1a9fcbf35b9aabf0e37b6c5eea296798)
5.13-rc1 [psi: Fix psi state corruption when schedule() races with cgroup move](https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/commit/?id=d583d360a620e6229422b3455d0be082b8255f5e)
| 2023/01/13 | Vincent Guittot <vincent.guittot@linaro.org> | [sched/fair: unlink misfit task from cpu overutilized](https://lore.kernel.org/all/20230113134056.257691-1-vincent.guittot@linaro.org) | TODO | v3 ☐☑✓ | [LORE](https://lore.kernel.org/all/20230113134056.257691-1-vincent.guittot@linaro.org) |
通过考虑 uclamp_min, task misfit 和 cpu overutilization 之间的 1:1 关系不再成立, 因为一个 util_avg 较小的任务可能由于 uclamp_min 的约束而不适合大容量的 cpu.
在 util_fits_cpu() 中添加一个新状态, 以反映任务适合 CPU 的情况, 除了 uclamp_min 提示 (这是一种性能要求).
使用 - 1 表示 CPU 不适合只是因为 uclamp_min, 因此我们可以使用这个新值采取额外的操作, 以选择不符合 uclamp_min 提示的最佳 CPU.
| 2023/01/12 | Daniel Bristot de Oliveira <bristot@kernel.org> | [sched/idle: Make idle poll dynamic per-cpu](https://lore.kernel.org/all/20230112162426.217522-1-bristot@kernel.org) | TODO | v1 ☐☑✓ | [LORE](https://lore.kernel.org/all/20230112162426.217522-1-bristot@kernel.org) |
| 2023/01/12 | Daniel Bristot de Oliveira <bristot@kernel.org> | [sched/idle: Make idle poll dynamic per-cpu](https://lore.kernel.org/all/20230112162426.217522-1-bristot@kernel.org) | TODO | v1 ☐☑✓ | [LORE](https://lore.kernel.org/all/20230112162426.217522-1-bristot@kernel.org) |
[鲲鹏 gcc mcmodel 选项详解](https://bbs.huaweicloud.com/blogs/272527)
[GCC for openEuler -mcmodel 选项详解](https://cdn.modb.pro/db/524836)
[对于几个锁的对比总结 Part1](https://blog.csdn.net/He11o_Liu/article/details/81077867)
[论文分享:Smartlocks: Lock Acquisition Scheduling for Self-Aware Synchronization](https://blog.csdn.net/He11o_Liu/article/details/81077695)
[论文分享 SANL:可扩展 NUMA-Aware 锁](https://blog.csdn.net/He11o_Liu/article/details/79255951)
[论文分享:Unlocking Energy](https://blog.csdn.net/He11o_Liu/article/details/81077777)
[论文分享:Non-scalable locks are dangerous](https://blog.csdn.net/He11o_Liu/article/details/80386839)
[转载 ---- 从 CPU cache 一致性的角度看 Linux spinlock 的不可伸缩性 (non-scalable)](https://blog.csdn.net/zhangshuaiisme/article/details/88147697)
[Scalable lock-free dynamic memory allocation 简要观感](https://blog.csdn.net/jollyjumper/article/details/53948391)
[从 CPU cache 一致性的角度看 Linux spinlock 的不可伸缩性 (non-scalable)](https://blog.csdn.net/dog250/article/details/80589442)
[[Paper 翻译]Scalable Lock-Free Dynamic Memory Allocation](https://blog.csdn.net/weixin_30457065/article/details/95622521)
[PV qspinlock 原理](https://blog.csdn.net/bemind1/article/details/118224344)
| 2023/01/26 | Waiman Long <longman@redhat.com> | [sched: Store restrict_cpus_allowed_ptr() call state](https://lore.kernel.org/all/20230127015527.466367-1-longman@redhat.com) | TODO | v3 ☐☑✓ | [LORE](https://lore.kernel.org/all/20230127015527.466367-1-longman@redhat.com) |
| 2023/01/20 | Wander Lairson Costa <wander@redhat.com> | [Fix put_task_struct() calls under PREEMPT_RT](https://lore.kernel.org/all/20230120150246.20797-1-wander@redhat.com) | TODO | v2 ☐☑✓ | [LORE v2,0/4](https://lore.kernel.org/all/20230120150246.20797-1-wander@redhat.com) |
| 2023/01/13 | Nathan Huckleberry <nhuck@google.com> | [workqueue: Add WQ_SCHED_FIFO](https://lore.kernel.org/all/20230113210703.62107-1-nhuck@google.com) | TODO | v1 ☐☑✓ | [LORE](https://lore.kernel.org/all/20230113210703.62107-1-nhuck@google.com) |
| 2018/11/11 | Paul E. McKenney <paulmck@linux.ibm.com> | [Automate initrd generation for v4.21/v5.0](https://lore.kernel.org/all/20181111200127.GA9511@linux.ibm.com) | 内核中引入 nolibc, 参见 LWN 报道 [Nolibc: a minimal C-library replacement shipped with the kernel](https://lwn.net/Articles/920158) | v5 ☐☑✓ | [LORE v5,0/8](https://lore.kernel.org/all/20181111200127.GA9511@linux.ibm.com) |
[McKenney: What Does It Mean To Be An RCU Implementation?](https://lwn.net/Articles/921351)
[GFP flags and the end of GFP_ATOMIC](https://lwn.net/Articles/920891)
[Linux Kernel Podcast](https://kernelpodcast.org)
[Reconsidering BPF ABI stability](https://lwn.net/Articles/921088)
| 2023/01/13 | Mel Gorman <mgorman@techsingularity.net> | [Discard `__GFP_ATOMIC`](https://lore.kernel.org/all/20230113111217.14134-1-mgorman@techsingularity.net) | TODO | v3 ☐☑✓ | [LORE v2,0/6](https://lore.kernel.org/all/20230109151631.24923-1-mgorman@techsingularity.net)<br>*-*-*-*-*-*-*-* <br>[LORE v3,0/6](https://lore.kernel.org/all/20230113111217.14134-1-mgorman@techsingularity.net) |
[Linux Developers Evaluating New "DOITM" Security Mitigation For Latest Intel CPUs](https://www.phoronix.com/review/intel-doitm-linux)
@@ -51,9 +51,26 @@ blogexcerpt: 虚拟化 & KVM 子系统
# 2 openEuler CONFIG_QOS_SCHED_SMT_EXPELLER
-------
openEuler-22.03 提供了内核驱离的特性, 通过 CONFIG_QOS_SCHED_SMT_EXPELLER 控制.
```cpp
8090ab77223b sched: Add tracepoint for qos smt expeller
42f42feeaae6 sched: Add statistics for qos smt expeller
fd5207be48fa sched: Implement the function of qos smt expeller
4e57e412b84a sched: Introduce qos smt expeller for co-location
```
[openEuler QOS_SCHED_SMT_EXPELLER 主体流程框架](./qos_smt_expeller.mmd)
# 3 OpenAnolis Group Identity 'Smt Expeller'
-------
| 日期 | 介绍 | 国际化 |
|:---:|:----:|:---:|
| 2021/02/08 | [Group Identity功能说明](https://www.alibabacloud.com/help/zh/elastic-compute-service/latest/group-identity-feature) | [Group identity feature](https://www.alibabacloud.com/help/en/elastic-compute-service/latest/group-identity-feature) |
<br>
@@ -41,8 +41,10 @@ flowchart TB
subgraph QoSSmtCheckNeedResched [执行驱逐]
direction TB
_qos_smt_check_need_resched --> for_each_siblings_cpu2;
%% 场景一: 当前核执行在线任务, 兄弟核执行离线任务, 需要进行驱逐.
%% 如果 THIS CPU 的 Siblings CPU 是 QOS_LEVEL_ONLINE, 但是 THIS CPU 正在执行的是 OFFLINE 任务, 需要进行驱逐, THIS CPU 触发 RESCHED, pick_next_task_fair 会选择 NULL.
for_each_siblings_cpu2 --Siblings CPU 上执行在线任务--> this_cpu_must_be_expeller("per_cpu(qos_smt_status, cpu) == QOS_LEVEL_ONLINE && task_group(current)->qos_level < QOS_LEVEL_ONLINE");
%% 场景二:当前核之前执行了驱逐,而兄弟 CPU 上从在线任务切换到离线任务,且当前无在线任务等待运行。
%% 如果 THIS CPU 的 Siblings CPU 是 QOS_LEVEL_OFFLINE, 且是 IDLE 状态, 而 THIS CPU 上只有离线任务, 没有在线任务, 同样需要触发 RESCHED, pick_next_task_fair 尝试选择一个 ONLINE 任务出来.
for_each_siblings_cpu2 --Siblings CPU 上执行离线任务--> this_cpu_can_run_IDLE("per_cpu(qos_smt_status, cpu) == QOS_LEVEL_OFFLINE && rq->curr == rq->idle && sched_idle_cpu(this_cpu)");
end
@@ -0,0 +1,22 @@
flowchart TB
__schedule --> pick_next_task["next = pick_next_task(rq, prev, &rf)"]
__schedule --> NotifySmtExpeller["notify_smt_expeller(rq, next)"]
subgraph NotifySmtExpeller ["保持 SMT 驱逐状态"]
direction TB
notify_smt_expeller --> DoNotifySmtExpeller;
subgraph DoNotifySmtExpeller ["保持 SMT 驱逐状态"]
direction TB
__notify_smt_expeller --> for_each_sibling_cpu["for_each_cpu(cpu, cpu_smt_mask(this_cpu)"] --> smp_send_reschedule["smp_send_reschedule(cpu)"];
end
end
scheduler_ipi --> handle_smt_expeller;
handle_smt_expeller --> __update_rq_on_expel;
handle_smt_expeller --> ExpelResched;
subgraph ExpelResched ["保持 SMT 驱逐状态"]
direction TB
handle_smt_expeller --> expel_resched --> set_tsk_need_resched;
end