Abstract:The frequent occurrence of hoisting accidents has caused great damage to the country, society and people. According to the video information in hoisting process, the accuracy and speed are the key to realize unmanned safety monitoring system. A new object segmentation network based on global coding and asymmetric convolution was proposed to study semi-supervised video object segmentation task. Firstly, video frames with labels were input into the network, and complementary features were extracted by global encoder and similarity encoder respectively, so that appearance of the target could be effectively represented. Then features of two branches were deeply fused through asymmetric convolution, and residual up-sampling decoding was used to generate prediction mask for target segmentation. The accuracy and overall indicator on DAVIS2017 dataset were 0.675 and 0.708 respectively, and frame rate was 31 frames per second. On hoisting dataset, the accuracy was 0.952, with the result of 0.976 overall indicator which was 5.1% higher than baseline, and frame rate was 26.16 frames per second. Compared to other methods on DAVIS2017 dataset and the hoisting dataset for experiment, the verification showed that the proposed method was effective in accuracy and speed.