融合几何特征的多模态点云分析网络
CSTR:
作者:
作者单位:

作者简介:

通讯作者:

中图分类号:

TP 391

基金项目:

国家自然科学基金资助项目(62273239); 上海市“科技创新行动计划”国内科技合作项目(20015801100)


Multimodal point cloud analysis network integrating geometric structural features
Author:
Affiliation:

Fund Project:

  • 摘要
  • |
  • 图/表
  • |
  • 访问统计
  • |
  • 参考文献
  • |
  • 相似文献
  • |
  • 引证文献
  • |
  • 资源附件
  • |
  • 文章评论
    摘要:

    多模态特征学习在点云领域得到广泛关注,但网络的分类分割精度仍很大程度上取决于点云分析骨干网络的设计。现有骨干网络在下采样阶段未能有侧重地压缩点云数量,在特征提取阶段难以高效捕捉多层次几何结构特征。针对上述问题,创新性地提出了一个融合多层次几何结构特征的点云分析网络,并引入多模态对齐预训练策略。在下采样阶段,将传统局部结构特征与全局注意力特征相结合,用以筛选出具有高判别力的结构关键点,从而大幅剔除对3D理解贡献有限的冗余点;在特征提取阶段,设计了一个轻量化模块,实现邻域中多层次几何结构特征的提取与自适应融合,在参数量与特征丰富性之间实现了良好的平衡;在预训练阶段,引入文本、图像、点云多模态特征对齐预训练策略,进一步增强模型的表现。实验结果表明,网络在ModelNet40和ScanObjectNN数据集上的分类精度分别达到了94.2%和87.9%,相较于经典和近期的点云分析网络均实现显著提升,在分割任务上也表现出了良好的性能。

    Abstract:

    Multimodal feature learning has received widespread attention in the point cloud domain, but the ultimate accuracy of the network in downstream tasks still depends heavily on the design of the point cloud analysis backbone network. Existing backbone networks fail to focus on compressing the number of point clouds in the down sampling phase, and it is difficult to efficiently capture multilevel geometric structure features in the feature extraction phase. Therefore, to address the above problems, a point cloud analysis network that incorporated multilevel geometric structure features was proposed, and a multimodal alignment pre-training strategy was introduced. In the down sampling stage, traditional local structural features were combined with global attention features to filter out structural key points with high discriminative power, thus substantially eliminating redundant points with limited contribution to 3D understanding. In the feature extraction stage, a lightweight module was designed to realize the extraction and adaptive fusion of multilevel geometric structural features in the neighboring domain, which achieved a good balance between the number of parameters and the richness of the features. In the pre-training stage, a multimodal feature alignment pre-training strategy for text, image, and point cloud was introduced to further enhance the performance of the model. The experimental results show that the network achieves classification accuracies of 94.2% and 87.9% on ModelNet40 and ScanObjectNN datasets, respectively, which is a significant improvement over both classical and recent point cloud analysis networks, and also shows good performance in segmentation tasks.

    参考文献
    相似文献
    引证文献
引用本文

郁森林,魏国亮,简晓富.融合几何特征的多模态点云分析网络[J].上海理工大学学报,2026,48(4):513-524.

复制
分享
相关视频

文章指标
  • 点击次数:
  • 下载次数:
  • HTML阅读次数:
  • 引用次数:
历史
  • 收稿日期:2025-06-13
  • 最后修改日期:
  • 录用日期:
  • 在线发布日期: 2026-09-15
  • 出版日期:
文章二维码