跳过正文
  1. Posts/

一次 AVX2 在 GCC 4 上 Core Dump 的排查经历

·1157 字·3 分钟

背景
#

起因是同事在实现 int4 功能的时候,流水线有一条死活过不了(gcc 版本为 4.8.5),一直 core dump。 经过初步排查,找出了如下最小复现代码:

 1
 2#include <immintrin.h>
 3
 4class Test{
 5    public:
 6    Test(){
 7        tmp = _mm256_set_epi32(0,0,0,0,0,0,0,0);
 8    }
 9    private:
10    __m256i tmp;
11};
12int main(){
13    auto *tmp = new Test();
14    return 0;
15}

gcc 版本为 4.8.5,编译选项为:

1g++ -std=c++11 -mavx2 a.cpp

现象为会 core 在 *tmp = _mm256_set_epi32(0,0,0,0,0,0,0,0); 这一行。

但是同样的代码、同样的编译选项,在 gcc 7.3 上就不会发生 core 的问题。

初步排查
#

查看汇编代码,gcc 4.8.5 生成的如下:

 1
 2main:
 3        push    rbp
 4        mov     rbp, rsp
 5        mov     edi, 32
 6        call    operator new(unsigned long)
 7        vpxor   xmm0, xmm0, xmm0
 8        vmovdqa YMMWORD PTR [rax], ymm0
 9        mov     eax, 0
10        pop     rbp
11        ret

链接在这里

然而在 gcc 7.3 下,生成的汇编代码如下:

 1
 2main:
 3        push    rbp
 4        mov     rbp, rsp
 5        push    r10
 6        sub     rsp, 8
 7        mov     esi, 32
 8        mov     edi, 32
 9        call    operator new(unsigned long, std::align_val_t)
10        vpxor   xmm0, xmm0, xmm0
11        vmovdqa YMMWORD PTR [rax], ymm0
12        mov     eax, 0
13        add     rsp, 8
14        pop     r10
15        pop     rbp
16        ret

链接在这里

发现调用的 new operator 竟然不是同一个:-std=c++17 下带了一个类型为 std::align_val_t 的参数。

同时观察到,如果不用 new 来创建 Object,也不会发生 core dump。

此时基本确定,问题和 new 有关。

new 的对齐规则
#

然后在公司大佬的指引下,看到了 -faligned-new

-faligned-new Enable support for C++17 new of types that require more alignment than void* ::operator new(std::size_t) provides. A numeric argument such as -faligned-new=32 can be used to specify how much alignment (in bytes) is provided by that function, but few users will need to override the default of alignof(std::max_align_t).

This flag is enabled by default for -std=c++17.

这个参数的作用其实是用来设置 __STDCPP_DEFAULT_NEW_ALIGNMENT__,这个值默认为 alignof(std::max_align_t)

可以用如下代码来验证:

 1#include <immintrin.h>
 2#include <iostream>
 3
 4class Test{
 5    public:
 6    Test(){
 7        tmp = _mm256_set_epi32(0,0,0,0,0,0,0,0);
 8    }
 9    private:
10    __m256i tmp;
11};
12int main(){
13    auto *tmp = new Test();
14    std::cout<<__STDCPP_DEFAULT_NEW_ALIGNMENT__;
15    return 0;
16}

编译选项为:

1g++ -std=c++17 -mavx2 -faligned-new=32 c.cpp

设置了有什么作用呢?

编译器会根据 __STDCPP_DEFAULT_NEW_ALIGNMENT__ 的值来判断调用哪个版本的 new。具体来说,如果 type 的 alignment 大于这个值,就会调用带对齐参数版本的 new:

1operator new(unsigned long, std::align_val_t)

否则就调用不带对齐参数版本的 new:

1operator new(unsigned long)

按照如上的推断,在 gcc 7.3、c++17 下,通过设置 -faligned-new 让编译器不去调用带对齐参数的 new,那么也应该发生 core 才对。

。。然而实际上并没有。使用如下编译参数,无事发生:

1g++ -std=c++17 -mavx2 -faligned-new=32 c.cpp

Why??? 为什么没有 core?

If you’re compiling in [c++17] mode only with a sufficiently recent compiler (e.g., GCC>=7, clang>=5, MSVC>=19.12), then everything is taken care by the compiler and you can stop reading.

因为 gcc 7 以后的编译器已经把对齐之类的事情帮我们做了。

按照这个想法,在 gcc 6 下,总会 core 吧?然后发现也没有。

继续排查发现,gcc 4.9.4 仍然会 core,但是 gcc 5 就没有问题了。

怀疑是 gcc 5 做了什么修复,或者是 gcc 5 对应的 glibc 做了什么修复。。不过暂时没有找到,这里待补充。

解决办法
#

手动对齐一下就好了

 1template <size_t ALIGNMENT>
 2struct alignas(ALIGNMENT) AlignedNew {
 3  static_assert(ALIGNMENT > 0, "ALIGNMENT must be positive");
 4  static_assert((ALIGNMENT & (ALIGNMENT - 1)) == 0,
 5      "ALIGNMENT must be a power of 2");
 6  static_assert((ALIGNMENT % sizeof(void*)) == 0,
 7      "ALIGNMENT must be a multiple of sizeof(void *)");
 8  static void* operator new(size_t count) { return Allocate(count); }
 9  static void* operator new[](size_t count) { return Allocate(count); }
10  static void operator delete(void* ptr) { free(ptr); }
11  static void operator delete[](void* ptr) { free(ptr); }
12
13 private:
14  static void* Allocate(size_t count) {
15    void* result = nullptr;
16    const auto alloc_failed = posix_memalign(&result, ALIGNMENT, count);
17    if (alloc_failed)  throw ::std::bad_alloc();
18    return result;
19  }
20};
21class Test: public AlignedNew<32> {
22    public:
23    Test(){
24        tmp = _mm256_set_epi32(0,0,0,0,0,0,0,0);
25    }
26    private:
27    __m256i tmp;
28};

参考链接
#

相关文章

tensorRT 模型兼容性说明

·525 字·2 分钟
名词说明 # CUDA. 一般来说指的是CUDA SDK. 目前经常使用的是CUDA 8.0和CUDA 10.1两个版本. 8.0和10.1都是SDK的版本号. CUDNN. The NVIDIA CUDA® Deep Neural Network library (cuDNN). 是一个可以为神经网络提供GPU加速的库 compute capability. 是GPU的固有参数,可以理解为GPU的版本.越新的显卡该数值往往越高. tensorRT.NVIDIA TensorRT™ is an SDK for high-performance deep learning inference. 是一个深度学习推理库,旨在提供高性能的推理速度. plan file,也称为 engine plan. 是生成的tensorRT 模型文件. 兼容性说明 # Engine plan 的兼容性依赖于GPU的compute capability 和 TensorRT 版本, 不依赖于CUDA和CUDNN版本.