首页 文章 精选 留言 我的

精选列表

搜索[accelerate],共198篇文章
优秀的个人博客,低调大师

BEVFormer-accelerate:基于EasyCV加速BEVFormer

作者:贺弘 夕陌 谦言 临在 导言 BEVFormer是一种纯视觉的自动驾驶感知算法,通过融合环视相机图像的空间和时序特征显式的生成具有强表征能力的BEV特征,并应用于下游3D检测、分割等任务,取得了SOTA的结果。我们在EasyCV开源框架(https://github.com/alibaba/EasyCV)中,对BEVFomer算法进行集成,并从训练速度、算法收敛速度角度对代码进行了一些优化。同时,我们进一步使用推理优化工具PAI-Blade对模型进行优化,相比于原始模型在A100配置下能取得40%的推理速度提升。本文将从以下几个部分进行介绍:1、BEVFormer算法思想 2、训练速度和算法收敛速度优化 3、使用PAI-Blade优化推理速度。 BEVFormer算法思想 如上图所示,BEVFormer由如下三个部分组成: backbone:用于从6个角度的环视图像中提取多尺度的multi-camera feature BEV encoder:该模块主要包括Temporal self-Attention 和 Spatial Cross-Attention两个部分。 Spatial Cross-Attention结合多个相机的内外参信息对对应位置的multi-camera feature进行query,从而在统一的BEV视角下将multi-camera feature进行融合。 Temporal self-Attention将History BEV feature和 current BEV feature通过 self-attention module进行融合。 通过上述两个模块,输出同时包含多视角和时序信息的BEV feature进一步用于下游3D检测和分割任务 Det&Seg Head:用于特定任务的task head BEVFormer训练优化 训加速优化 我们从数据读取和减少内存拷贝消耗等角度对训练代码进行优化。 数据读取 使用更高效的图片解码库 turbojpeg BEVFormer在训练过程中,需要时序上的数据作为输入,将串形的读取方式优化为并行读取。 先做resize再做其他预处理,减少了额外像素带来的计算开销 内存拷贝优化 使用pin_memery=True,并修复了mmcv DataContainer pin_memory的bug 将代码中的numpy操作替换为torch.tensor,避免不必要的h2d拷贝 other 使用torch.backends.cudnn.benchmark=True(ps:需要保证在输入数据没有动态性的情况下使用,否则反而会增加训练耗时) 修复了torch.cuda.amp混合精度在LayerNorm层失效的bug 我们在A100 80G的机器上,使用fp16对比吞吐量如下: Setting throughput(samples/s) BEVFormer-tiny bs=32 3.55 EasyCV BEVFormer-tiny bs=32 9.84( +177% ) BEVFormer-base bs=5 0.727 EasyCV BEVFormer-base bs=5 0.8( +10% ) 精度收敛优化 我们使用额外的数据增广方式和不同的损失函数来优化模型。同时加入额外的训练策略来进一步提升模型收敛速度及精度。 数据增广方式 rand scale(采用不同分辨率的输入进行训练,实验中发现该操作会引入至少20%的额外训练时间,因此在下述实验中,均没有采用) rand_flip(以50%的概率随机翻转图片) 损失函数 使用smooth l1 loss或 balance l1 loss代替l1 loss。(在mini dataset的实验中,这两个损失都可以提升精度,下面的实验中采用balance l1 loss) 训练策略 使用one2many Branch 这个做法来自于H-Deformable-DETR,在DETR系列的检测模型中采用one2one的匹配方式来分配GT Boxes,这种做法虽然让模型在测试的时候,能够避免冗余的NMS后处理操作,但是只有少数的Query会被分配给正样本,导致训练时模型收敛速度相比于one2many的方式会慢很多。因此,在训练过程中加入auxiliary Query,同一个GT Box会匹配多个auxiliary Query,并使用attention mask将one2one branch和one2many branch的信息隔离开。通过这样的方式,能够显著的提升训练过程中的收敛速度,同时在测试过程中只需要保持one2one branch进行预测。(在实验中,使用额外加入1800个auxiliary Query,每个GT box匹配4个query进行训练) CBGS in one2many Branch 我们的实验是在NuScenes数据集上进行的,在该数据集的3D检测任务上有10类标签,但是这10类标签之间的样本极度不均衡,很多算法会采用CBGS操作进行类间样本均衡,但是这个操作会将整个数据集扩大4.5倍,虽然有一定的精度提升,但是也带来了巨大的训练成本。我们考虑在one2many Branch上进行样本均衡操作,即对于实例数量较多的样本使用较少的auxiliary Query进行匹配,而对于长尾的样本使用较多的auxiliary Query进行匹配。通过CBGS in one2many Branch的方式,训练时间和base保持一致的基础上会进一步提升收敛速度,最终的精度也有一定的提升。(实验中匹配框数量变化:[4, 4, 4, 4, 4, 4, 4, 4, 4, 4] -> [2, 3, 7, 7, 9, 6, 7, 6, 2, 5]) 我们在单机8卡A100 80G下进行实验,如下表所示: config setting NDS mAP throughput(samples/s) 官方 BEVFormer-base 52.44 41.91 3.289 EasyCV BEVFormer-base 52.66 42.13 3.45 EasyCV BEVFormer-base-one2manybranch 53.02(+0.58) 42.48(+0.57) 3.40 EasyCV BEVFormer-base-cbgs_one2manybranch 53.28(+0.84) 42.63(+0.72) 3.41 模型收敛速度如下图所示: 由上图可以看出,使用上述优化方式可以大幅提升模型收敛速度,仅需要75%的训练时间就可以达到base的最终精度。同时最终的NDS相比于base也有0.8的提升。 详细配置,训练log和模型权重,参考:https://github.com/alibaba/EasyCV/blob/master/docs/source/model_zoo_det3d.md 在阿里云机器学习平台PAI上使用BEVFormer模型 PAI-DSW(Data Science Workshop)是阿里云机器学习平台PAI开发的云上IDE,面向各类开发者,提供了交互式的编程环境。在DSW Gallery中(链接),提供了各种Notebook示例,方便用户轻松上手DSW,搭建各种机器学习应用。我们也在DSW Gallery中上架了BEVFormer进行3D检测的Sample Notebook(见下图),欢迎大家体验! 使用PAI-Blade进行推理加速 PAI-Blade是由阿里云机器学习平台PAI开发的模型优化工具,可以针对不同的设备不同模型进行推理加速优化。PAI-Blade遵循易用性,鲁棒性和高性能为原则,将模型的部署优化进行高度封装,设计了统一简单的API,在完成Blade环境安装后,用户可以在不了解ONNX、TensorRT、编译优化等技术细节的条件下,通过简单的代码调用方便的实现对模型的高性能部署。更多PAI-Blade相关技术介绍可以参考 [PAI-Blade介绍]。 PAI-EasyCV中对Blade进行了支持,用户可以通过PAI-EasyCV的训练config 中配置相关export 参数,从而对训练得到的模型进行导出。 对于BEVFormer模型,我们在A100机器下进行进行推理速度对比,使用PAI-Blade优化后的模型能取得42% 的优化加速。 Name Backend Median(FPS) Mean(FPS) Median(ms) Mean(ms) easycv TensorRT 3.68697 3.68651 0.271226 0.271259 easycv script TensorRT 3.8131 3.79859 0.262254 0.26337 blade TensorRT 5.40248 5.23383**(+42%)** 0.1851 0.192212 环境准备 我们提供一个PAI-Blade + PAI-EasyCV 的镜像包供用户可以直接使用,镜像包地址:easycv-blade-torch181-cuda111.tar 用户也可以基于Blade每日发布的镜像自行搭建推理环境 [PAI-Blade社区镜像发布]。 自行搭建环境时需要注意:BEVFomer-base使用resnet101-dcn作为image backbone,DCN算子使用的是mmcv中的自定义算子,为了导出TorchScript,我们对该接口进行了修改。所以mmcv需要源码编译。 clone mmcv源码 $ git clone https://github.com/open-mmlab/mmcv.git 替换mmcv文件 替换时请注意mmcv的版本,注意接口要匹配。mmcv1.6.0版本已验证。 参考easycv/thirdparty/mmcv/目录下的修改文件。用mmcv/ops/csrc/pytorch/modulated_deform_conv.cpp和mmcv/ops/modulated_deform_conv.py去替换mmcv中的原文件。 源码编译 mmcv源码编译请参考:https://mmcv.readthedocs.io/en/latest/get_started/build.html 导出Blade模型 导出Blade的模型的配置可以参考文件bevformer_base_r101_dcn_nuscenes.py中的export字段,配置如下: export = dict( type='blade', blade_config=dict( enable_fp16=True, fp16_fallback_op_ratio=0.0, customize_op_black_list=[ 'aten::select', 'aten::index', 'aten::slice', 'aten::view', 'aten::upsample', 'aten::clamp' ] ) ) 导出命令: $ cd ${EASYCV_ROOT} $ export PYTHONPATH='./' $ python tools/export.py configs/detection3d/bevformer/bevformer_base_r101_dcn_nuscenes.py bevformer_base.pth bevformer_export.pth Blade模型推理 推理脚本: from easycv.predictors import BEVFormerPredictor blade_model_path = 'bevformer_export.pth.blade' config_file = 'configs/detection3d/bevformer/bevformer_base_r101_dcn_nuscenes.py' predictor = BEVFormerPredictor( model_path=blade_model_path, config_file=config_file, model_type='blade', ) inputs_file = 'nuscenes_infos_temporal_val.pkl' # 以NuScenes val数据集文件为例 input_samples = mmcv.load(inputs_file)['infos'] predict_results = predictor(input_samples) print(predict_results) NuScenes数据集准备请参考:NuScenes数据集准备 展望 我们在EasyCV框架中,集成了BEVFormer算法,并从训练加速、精度收敛和推理加速角度对算法进行了一些改进。近期,也涌现了许多新的BEV感知算法,如BEVFormerv2。在BEVFormerv2中通过Perspective Supervision的方式,让算法能够不受限于使用一些在深度估计或3D检测上的预训练backbone,而直接使用近期更有效的大模型BackBone(如ConvNext、DCNv3等),同时采用two-stage的检测方式进一步增强模型能力,在Nuscenes数据集的camera-based 3D检测任务取得sota的结果。 EasyCV(https://github.com/alibaba/EasyCV)会持续跟进业界sota方法,欢迎大家关注和使用,欢迎大家各种维度的反馈和改进建议以及技术讨论,同时我们十分欢迎和期待对开源社区建设感兴趣的同行一起参与共建。 EasyCV往期分享 EasyCV开源地址:https://github.com/alibaba/EasyCV 使用 EasyCV Mask2Former轻松实现图像分割 https://zhuanlan.zhihu.com/p/585259059 EasyCV DataHub 提供多领域视觉数据集下载,助力模型生产 https://zhuanlan.zhihu.com/p/572593950 EasyCV带你复现更好更快的自监督算法-FastConvMAE https://zhuanlan.zhihu.com/p/566988235 基于EasyCV复现DETR和DAB-DETR,Object Query的正确打开方式 https://zhuanlan.zhihu.com/p/543129581 基于EasyCV复现ViTDet:单层特征超越FPN https://zhuanlan.zhihu.com/p/528733299 MAE自监督算法介绍和基于EasyCV的复现 https://zhuanlan.zhihu.com/p/515859470 EasyCV开源|开箱即用的视觉自监督+Transformer算法库 https://zhuanlan.zhihu.com/p/50521999

优秀的个人博客,低调大师

How to Accelerate Your Python Deep Learning with Cloud GPU?

Overloaded This afternoon, I trained a 3-layers neural network as a regression model to predict the house price in Boston district with Python and Keras. The example case came from the book "Deep Learning with Python". There were 2 big loop during the running procedure. The first one went through the data for 100 times (epochs), while the second one ran 500 epochs. My poor laptop was apparently overladed in such a hot summer weather and the fan was roaring. It seems the laptop is not the best choice to train deep neural models. It would be so great if I have got a GPU. Suddenly, it occurs to me that it is not necessary to train the model locally. It's a cloud computing age! How about to run the code on cloud GPU to save my laptop's effort? Encounter It reminds me a video clip post by Siraj Raval on Youtube recently. He recommended cloud GPU platform, namely Floydhub, in this video. Actually, I once tried AWS GPU product in a online deep learning course. The instructor collaborated with AWS and provided all the students with AWS Computing power to solve the exercise as well as the homework. However, it was not a very good experience, since he had to make a long video to show the students how to configure the AWS instance. Indeed, comparing with some other solutions, the AWS was simple enough, yet still not so simple for the new newbies. The website FloydHub, on the other hand, solved the pain point well. Firstly, it is wrapper over AWS, and filtered out a lot of complex operations. Secondly, FloydHub is batteries-included with a lot of main stream machine learning frameworks. Besides, it is well-documented and friendly to the new users. The slogan is: Focus on what matters. Let FloydHub handle the grunt work. Honestly, I like all the things designed for the lazy folks. So I registered immediately and validated my email. Then I got 2 hours GPU running time for free! To spend the precious GPU running time on something import, I read the Quick Start Tutorial eagerly. Several minutes later, I feel confident to use it. Trial I created a new job from personal control panel on FloydHub and named it "try-keras-boston-house-regression". Then I exported a Python Script file from my local Jupyter Notebook. I created a new directory and copied the script file into it. To save the Evaluation Metrics of the training and evaluation process, I added 3 lines of code in the end of the Python Script. import pickle with open('data.pickle', 'wb') as f: pickle.dump([all_scores, all_mae_histories], f) In this way, we can save all_scores and all_mae_histories data into a file named data.pickle with the Pickle Module in Python. Then let's dive into the shell and navigate to this new created folder with cd command and execute the following command: pip install floyd-cli The command line interface of FloydHub is ready to use. We can login the FloydHub account with: floyd login Then input your FloydHub username and password. When it's ready, run: floyd init try-keras-boston-house-regression Please notice the last parameter should be identical to the title you input just now when created the new job from control panel. Now we can run the Python script with following command: floyd run --gpu --env tensorflow-1.8 "python 03-house-price.py" In this command, --gpu means that we ask the FloydHub to run the script in a GPU environment instead of a default CPU one, and --env tensorflow-1.8 means it will use Tensorflow version 1.8, and the Keras version is 2.1.6 accordingly. If you want to use other framework or choose a different version, please refer to this link. In response, we get the following messages from FloydHub. It's all set. Yes, so easy. And your learning job is already running in the cloud. Results While the job was running, I drank some tea, read several pages of books and browsed some news on Social Media with my phone. When the running job is done, it will terminate the environment and will not charge you any extra GPU running time. So you don't need to keep an eye on it. When I came back to my computer, the job's already fininished. GPU memory was busy during the whole procedure, as the Utilization was above 90% most of the time. The GPU, on the other hand, was not busy at all. Maybe my neural network was too simple. Scrolling down the page, we can see the logs. The output was similar to the one when you train the model locally. Besides, it showed you extra information about GPU resource allocation. To see the saved file, you can open the Files tag. The pickle file's already there. FloydHub helped us with all the hard computing job, and my laptop is much cooler this time. You can download the pickle file, and put it back into the original working directory. Let's go back to the Jupyter Lab page on the laptop and open a new ipynb file. The following code can check the running results. import pickle import matplotlib.pyplot as plt import numpy as np %matplotlib inline with open('data.pickle', 'rb') as f: [all_scores, all_mae_histories] = pickle.load(f) num_epochs = 500 average_mae_history = [ np.mean([x[i] for x in all_mae_histories]) for i in range(num_epochs) ] plt.plot(range(1, len(average_mae_history) + 1), average_mae_history) plt.xlabel('Epochs') plt.ylabel('Validation MAE') plt.show() Please notice these codes will only do some drawings. Here is the result: The visualization result is identical to the textbook which shows the code ran smoothly on the Cloud GPU environment. You can check the remaining GPU running time easily. There's still more than 1 hour to play with. Great! Workspace Just now, I showed you how to run FloydHub in Command Line Interface. If you are familiar with bash command, it will be great. However, for the new users who do not want to use the shell command, I recommend you to try an easier way. Click the Workspace tab. You will see two existing Workspace examples. Try to open the first one and check it out. Hit the green Resume button on top right, the system will try to provide us the environment. When it's done, you'll see the familiar Jupyter lab interface. Open the dog-breed-classification.ipynb from the left side file list. It's a complete example to separate different dog breeds. Hit Run -> Restart Kernel and Run All Cells from the menu. You'll figure out there is no significant difference with running the code locally. However, this time, you are using GPU! What if you want to set up a new workspace yourself? You can go back to the Project page . For each project, you can create new workspace with the Create Workspace button. Floydhub will ask you how to create the new workspace. Let's select Start from scratch on the left side and choose the environment. Let's change the default one into Tensorflow 1.9 and GPU. Hit the Create Workspace. Then click on the link try-keras-boston-house-regression workspace. A Jupyter Lab interface is ready. You don't need to install Tensorflow or configure the GPU yourself. Even better, you don't need to run bash commands this time. Just input the Python code, and use Keras and Tensorflow freely. That's cool! Start your own Deep Learning Journey with Floydhub. Summary You don't need to buy your own expensive deep learning device if you just need GPU computing power occasionally. It will be a waste, and you'll not get a good price when you want to sell it to make an upgrade. In this case, Cloud GPU is a better choice. Have you ever used any other Cloud GPUs? What are the pros and cons comparing with Floydhub? I would like to have your feedbacks.

资源下载

更多资源
Mario

Mario

马里奥是站在游戏界顶峰的超人气多面角色。马里奥靠吃蘑菇成长,特征是大鼻子、头戴帽子、身穿背带裤,还留着胡子。与他的双胞胎兄弟路易基一起,长年担任任天堂的招牌角色。

Spring

Spring

Spring框架(Spring Framework)是由Rod Johnson于2002年提出的开源Java企业级应用框架,旨在通过使用JavaBean替代传统EJB实现方式降低企业级编程开发的复杂性。该框架基于简单性、可测试性和松耦合性设计理念,提供核心容器、应用上下文、数据访问集成等模块,支持整合Hibernate、Struts等第三方框架,其适用范围不仅限于服务器端开发,绝大多数Java应用均可从中受益。

Sublime Text

Sublime Text

Sublime Text具有漂亮的用户界面和强大的功能,例如代码缩略图,Python的插件,代码段等。还可自定义键绑定,菜单和工具栏。Sublime Text 的主要功能包括:拼写检查,书签,完整的 Python API , Goto 功能,即时项目切换,多选择,多窗口等等。Sublime Text 是一个跨平台的编辑器,同时支持Windows、Linux、Mac OS X等操作系统。

WebStorm

WebStorm

WebStorm 是jetbrains公司旗下一款JavaScript 开发工具。目前已经被广大中国JS开发者誉为“Web前端开发神器”、“最强大的HTML5编辑器”、“最智能的JavaScript IDE”等。与IntelliJ IDEA同源,继承了IntelliJ IDEA强大的JS部分的功能。

用户登录
用户注册