跳过正文
  1. Posts/

Jetson Nano踩坑记录

·3101 字·7 分钟

写在前面
#

主要是需要在 jetson nano 上做模型转换,来记录下踩的坑。 目前有两条路径,一条是我们现有的转换路径,也就是 pytorch->onnx(->caffe)->trt 的路径。 在这条路径上踩了比较多的坑,最终暂时放弃,最直接的原因是 cudnn8.0 升级接口发生改动,编译 caffe 遇到较多问题。 这里其实仍然采用了两条平行的路径,一条是直接在 nano 上构建环境,另外一种是基于 docker(包括构建交叉编译环境用于加快编译速度)。

另一条路径是基于 torch2trt,是一条直接 pytorch->trt 的路径。 这里主要记录在第一条路径上踩过的坑。

环境准备
#

先过一遍开发者手册 主要是介绍了下 nano 的硬件和 jetpack 的组件, jetpack 可以理解成一个 nvidia 用在嵌入式设备上的 SDK 工具包。

镜像烧录
#

然后需要烧录镜像,参考Write Image to the microSD Card 最初使用 Etcher 来烧录,没有启动起来,原因未知。使用终端命令烧录就没问题了。

更换国内 apt 源
#

我拿到的 jetson nano 是基于 ubuntu 18.04 LTS 的版本,默认源比较慢,换成国内源。需要注意要换 arm 版本的源。

1wget -O /etc/apt/sources.list https://repo.huaweicloud.com/repository/conf/Ubuntu-Ports-bionic.list
2apt update

基于docker的模型转换
#

由于在x86上模型转换的工具是基于docker的,因此最初的想法是在nano上也基于docker进行转模型,这样或许可以复用一些x86上的工作。

确认了 Jetson Nano 上是可以跑 docker 和 NVIDIA Container Runtime on Jetson (Beta) 的。 但是最初发现 arm 的 cuda docker image 最低支持 cuda11.0 的版本: cuda-arm64 docker image tag。 然而目前(2020 年 9 月)最新的 jetpack4.4 版本只支持到 cuda10.2 的版本。 尝试了下使用 cuda-arm64 的 docker image,发现驱动版本不够,而 nano 上似乎很难升级驱动版本,因此差点就放弃这条路径了。

然而发现,在 nano 上要用的 cuda 镜像并不应该是 cuda-arm,而是 NVIDIA L4T Base

结论
#

参考 Jetson 开发者手册归档,发现一个可能的问题是,JetPack 中 cuDNN、TensorRT 等库的版本不是连续的。

  • jetpack4.4:CUDA 10.2、TensorRT 7.1.3、cuDNN 8.0.0
  • jetpack4.3:CUDA 10.0.326、TensorRT 6.0.1.10、cuDNN 7.6.3
  • jetpack4.2.1:CUDA 10.0.326、TensorRT 5.1.6.1、cuDNN 7.5.0.56
  • jetpack4.2.0:CUDA 10.0.166、TensorRT 5.0.6.3、cuDNN 7.3.1.28

发现并不能知道 TensorRT7.0 的版本,arm 版本的库似乎官方也没有提供地址下载。 最初想着是将 cuDNN 7.6.3 从 jetpack4.3 的镜像中拿出来,拷贝过去使用, 然后发现 TensorRT 7.1.3 是依赖 cuDNN8.0 的……死了。或许可以考虑把工具链中的 caffe 摘除掉。

编译 opencv
#

这部分其实发生在上部分得到结论之前, 先尝试下在这个镜像上编译 opencv。 提示找不到 libjasper-dev,似乎是个可选的包,先不装试试看。 然后提示 fatal error: LAPACKE_H_PATH-NOTFOUND, 解决办法参考 opencv install

  • install package liblapacke-dev
  • manually define -DLAPACKE_H_PATH=/usr/include when calling cmake

由于在 nano 上编译速度非常感人,因此在 x86 上配置了 arm 的交叉编译环境,所以实际上是在 x86 机器上的交叉编译环境中进行编译的。

安装caffe2TRT工具链
#

cmake 版本 3.10 太低,需要升级,编译 3.16 版本 cmake 的时候报错: Could NOT find OpenSSL 通过

1
2sudo apt-get install openssl
3
4sudo apt-get install libssl-dev

解决

然后需要装 libgoogle-glog-dev 和 libgflags-dev。

报错: unrecognized command line option ‘-mavx2’ 原因是这个是 x86 的编译选项,arm 平台不支持,解决办法是在 CMakeLists 中去掉这个编译选项。

编译 caffe 报错: fatal error: boost/random/mersenne_twister.hpp: No such file or directory #include “boost/random/mersenne_twister.hpp” 应该是 boost 相关的库没装, apt install libboost-all-dev 之后解决。

然后报错 error: ‘iota’ is not a member of ‘std’ 解决办法是头文件添加了 #include

需要装 git lfs,试了下 apt install git-lfs,可以解决。

发现 nano 的 TensorRT、cuDNN 头文件在 /usr/include/aarch64-linux-gnu, 动态库在 /usr/lib/aarch64-linux-gnu。

编译 caffe 报错: error: identifier “cublasStatus_t” is undefined 这个符号应该是定义在 /usr/local/cuda/include/cublas.h 中的。 然后发现 nano 上 cuda10.2 版本,cublas 相应的头文件是在 /usr/include 的, 相应的动态库是在 /usr/lib/aarch64-linux-gnu 下的。 去 /usr/include 下发现 cublas 相关的 header file 大小竟然都是 0……??? 于是去 nano 上把 cublas 相关的头文件拷贝过来,问题解决。

继续报错: error: ‘iota’ is not a member of ‘std’ std::iota(idxs.begin(), idxs.end(), 0);

解决办法: 添加 #include

接着报错: error: ‘CUDNN_MAJOR’ was not declared in this scope 应该是编译的时候没找到 cudnn 导致的。 发现是 cudnn8 的头文件里没有了 cudnn.h,caffe 会去这个文件中找版本, 于是加了个软链接解决

1ln -s cudnn_version_v8.h  cudnn.h

接下来报错: caffe/util/cudnn.hpp(20): error: identifier “cudnnStatus_t” is undefined make VERBOSE=1 观察,是已经有了头文件所在的 /usr/include 路径的。

发现是 cudnn8 里面,头文件不止需要一个 cudnn.h,拆成了好多个头文件, 于是在 cudnn.h 里添加了

1#include <cudnn_ops_infer_v8.h>

解决

接下来报错 error: identifier “cudnnConvolutionDescriptor_t” is undefined 在cudnn.h里添加了

1#include <cudnn_cnn_infer_v8.h>

解决

要改的内容太多了,工作量不可控,放弃进一步的修改,转而采用的思路是移除掉转换工具对 caffe 的依赖。

直接在裸机上安装环境
#

安装 pytorch
#

注意由于架构不同,不能直接 pip install 来安装 pytorch。 参考 pytorch-for-jetson-version-1-6-0-now-available

先尝试安装 pip3,用在 x86 上的方式没成功。 最后是使用 https://pip.pypa.io/en/stable/installing/ 中的 get-pip.py 来安装。

于是尝试 pip3 install torch1.5.whl, 安装正常结束,但是

1python -c 'import torch'

的时候报错: “OSError: libmpi_cxx.so.20: cannot open shared object file: No such file or directory”

1apt install libopenmpi2

后解决。 接下来报错: ImportError: libopenblas.so.0: cannot open shared object file: No such file or directory

1apt install libopenblas-dev

后解决。

再之后直接 segment fault 了。 参考 在Jetson TX2上安装Python,JetsonTX2,傻瓜式,pytorch 等帖子,似乎是 arm 版本的 pytorch 1.5 安装包有问题。 尝试了 pytorch 1.4 和 pytorch 1.6,果然都没问题,有点坑。 最后决定使用 pytorch 1.6 版本。

安装 onnx
#

参考 build-onnx-on-arm-64 执行了

1pip3 install cython protobuf numpy
2sudo apt-get install libprotobuf-dev protobuf-compiler
3pip3 install onnx

然后在编译 onnx 的时候报错: “onnx/third_party/pybind11/include/pybind11/detail/common.h:112:10: fatal error: Python.h: No such file or directory” 参考 Make Error, fatal error: Python.h: No such file or directory compilation terminated,在执行

1apt install python3-dev

后解决

之后在 import caffe_pb2 的时候,报错: " options=_descriptor._ParseOptions(descriptor_pb2.FieldOptions(), _b(’\020\001’)), file=DESCRIPTOR), TypeError: new() got an unexpected keyword argument ‘file’" 原因是 protobuf 版本 3.0.0 太低了,file 这个参数是 protobuf 3.5 之后的版本引入的,于是升级 protobuf 版本到 3.5.1:

1pip install protobuf==3.5.1

问题解决

在 x86 上交叉编译
#

构建交叉编译的镜像
#

发现 nano 上编译速度过于感人,因此尝试使用交叉编译的方式。 参考 Enabling Jetson Containers on an x86 workstation (using qemu)

结果如下:

 1$ sudo apt-get install qemu binfmt-support qemu-user-static
 2# Check if the entries look good.
 3$ sudo cat /proc/sys/fs/binfmt_misc/status
 4enabled
 5
 6# See if /usr/bin/qemu-aarch64-static exists as one of the interpreters.
 7$ cat /proc/sys/fs/binfmt_misc/qemu-aarch64
 8enabled
 9interpreter /usr/bin/qemu-aarch64-static
10flags: OC
11offset 0
12magic 7f454c460201010000000000000000000200b700
13mask ffffffffffffff00fffffffffffffffffeffffff

注意少了个 ‘F’ flag。

Make sure the F flag is present, if not head to the troubleshooting section, as this will result in a failure to start the Jetson container.

参考 Running or building a container on x86 (using qemu+binfmt_misc) is failing 执行

1
2 docker run --rm --privileged multiarch/qemu-user-static --reset -p yes -c yes

喜提一屏幕报错:

详细代码
 1
 2 Setting /usr/bin/qemu-alpha-static as binfmt interpreter for alpha
 3sh: write error: Invalid argument
 4Setting /usr/bin/qemu-arm-static as binfmt interpreter for arm
 5sh: write error: Invalid argument
 6Setting /usr/bin/qemu-armeb-static as binfmt interpreter for armeb
 7sh: write error: Invalid argument
 8sh: write error: Invalid argument
 9Setting /usr/bin/qemu-sparc-static as binfmt interpreter for sparc
10Setting /usr/bin/qemu-sparc32plus-static as binfmt interpreter for sparc32plus
11sh: write error: Invalid argument
12Setting /usr/bin/qemu-sparc64-static as binfmt interpreter for sparc64
13sh: write error: Invalid argument
14Setting /usr/bin/qemu-ppc-static as binfmt interpreter for ppc
15sh: write error: Invalid argument
16Setting /usr/bin/qemu-ppc64-static as binfmt interpreter for ppc64
17sh: write error: Invalid argument
18Setting /usr/bin/qemu-ppc64le-static as binfmt interpreter for ppc64le
19sh: write error: Invalid argument
20Setting /usr/bin/qemu-m68k-static as binfmt interpreter for m68k
21sh: write error: Invalid argument
22sh: write error: Invalid argument
23Setting /usr/bin/qemu-mips-static as binfmt interpreter for mips
24Setting /usr/bin/qemu-mipsel-static as binfmt interpreter for mipsel
25sh: write error: Invalid argument
26Setting /usr/bin/qemu-mipsn32-static as binfmt interpreter for mipsn32
27sh: write error: Invalid argument
28Setting /usr/bin/qemu-mipsn32el-static as binfmt interpreter for mipsn32el
29sh: write error: Invalid argument
30sh: write error: Invalid argument
31Setting /usr/bin/qemu-mips64-static as binfmt interpreter for mips64
32Setting /usr/bin/qemu-mips64el-static as binfmt interpreter for mips64el
33sh: write error: Invalid argument
34Setting /usr/bin/qemu-sh4-static as binfmt interpreter for sh4
35sh: write error: Invalid argument
36Setting /usr/bin/qemu-sh4eb-static as binfmt interpreter for sh4eb
37sh: write error: Invalid argument
38Setting /usr/bin/qemu-s390x-static as binfmt interpreter for s390x
39sh: write error: Invalid argument
40Setting /usr/bin/qemu-aarch64-static as binfmt interpreter for aarch64
41sh: write error: Invalid argument
42Setting /usr/bin/qemu-aarch64_be-static as binfmt interpreter for aarch64_be
43sh: write error: Invalid argument
44sh: write error: Invalid argument
45Setting /usr/bin/qemu-hppa-static as binfmt interpreter for hppa
46Setting /usr/bin/qemu-riscv32-static as binfmt interpreter for riscv32
47sh: write error: Invalid argument
48Setting /usr/bin/qemu-riscv64-static as binfmt interpreter for riscv64
49sh: write error: Invalid argument
50Setting /usr/bin/qemu-xtensa-static as binfmt interpreter for xtensa
51sh: write error: Invalid argument
52Setting /usr/bin/qemu-xtensaeb-static as binfmt interpreter for xtensaeb
53sh: write error: Invalid argument
54Setting /usr/bin/qemu-microblaze-static as binfmt interpreter for microblaze
55sh: write error: Invalid argument
56Setting /usr/bin/qemu-microblazeel-static as binfmt interpreter for microblazeel
57sh: write error: Invalid argument
58Setting /usr/bin/qemu-or1k-static as binfmt interpreter for or1k
59sh: write error: Invalid argument

查到 sh: write error: Invalid argument - Centos 7

发现原因是 “-p yes” 这个参数是在 4.10 版本的 kernel 后才支持的,而我本地的机器 ubuntu 14.04 的 kernel 版本是 4.2.0-27-generic。

绕过去的办法是从 qemu-user-static 的 docker 里 copy 一些必要的文件出来:

1$ docker build --rm -t my-aarch64-ubuntu -<<EOF
2FROM multiarch/qemu-user-static:x86_64-aarch64 as qemu
3FROM arm64v8/ubuntu
4COPY --from=qemu /usr/bin/qemu-aarch64-static /usr/bin
5EOF
6
7$ docker run --rm -t my-aarch64-ubuntu uname -m
8aarch64

接下来可以参考 Running and Building ARM Docker Containers on x86

编译 cuda 有关的库
#

然而上面这些步骤只是能够编译和 cuda 无关的部分。 发现对于相同的镜像,在 x86 上启动时 /usr/local/cuda/lib64 会少很多内容,但是在 nano 上启动就是正常的, 怀疑是 nano 的 nvidia-docker 做了什么事情。

参考 cross-compile-cuda-engines 以及 [l4t-base] Building CUDA code in a Jetson Container on an x86 workstation, 发现一个 workaround 是把 nano 上 cuda 里面的动态库 copy 到镜像里面。

参考链接
#

相关文章