Abstract:Multimodal feature learning has received widespread attention in the point cloud domain, but the ultimate accuracy of the network in downstream tasks still depends heavily on the design of the point cloud analysis backbone network. Existing backbone networks fail to focus on compressing the number of point clouds in the down sampling phase, and it is difficult to efficiently capture multilevel geometric structure features in the feature extraction phase. Therefore, to address the above problems, a point cloud analysis network that incorporated multilevel geometric structure features was proposed, and a multimodal alignment pre-training strategy was introduced. In the down sampling stage, traditional local structural features were combined with global attention features to filter out structural key points with high discriminative power, thus substantially eliminating redundant points with limited contribution to 3D understanding. In the feature extraction stage, a lightweight module was designed to realize the extraction and adaptive fusion of multilevel geometric structural features in the neighboring domain, which achieved a good balance between the number of parameters and the richness of the features. In the pre-training stage, a multimodal feature alignment pre-training strategy for text, image, and point cloud was introduced to further enhance the performance of the model. The experimental results show that the network achieves classification accuracies of 94.2% and 87.9% on ModelNet40 and ScanObjectNN datasets, respectively, which is a significant improvement over both classical and recent point cloud analysis networks, and also shows good performance in segmentation tasks.