← 返回文章列表

2026 · 08 · 04   /   技术分享

技术分享

100 篇值得长期收藏的论文与工程资料,另附 GPU、CUDA、内存和存储的硬件观察栏目。

论文优先指向原文或会议页面;工程与硬件资料优先指向官方文档。链接会随项目演进而更新,适合作为学习路线和采购前核对的起点,而非唯一答案。

55机器学习与深度学习论文
15C++ 开发资料
15Rust 开发资料
15数据库与分布式系统资料
持续更新算力与硬件观察

机器学习与深度学习论文

从经典统计学习、树模型和降维,到视觉、生成式模型、语言模型与强化学习。

  1. Support-Vector Networks支持向量机(SVM)
  2. Random Forests随机森林
  3. A Decision-Theoretic Generalization of On-Line Learning and an Application to BoostingAdaBoost
  4. Greedy Function Approximation: A Gradient Boosting Machine梯度提升树
  5. XGBoost: A Scalable Tree Boosting SystemXGBoost
  6. LightGBM: A Highly Efficient Gradient Boosting Decision TreeLightGBM
  7. CatBoost: unbiased boosting with categorical featuresCatBoost
  8. A Density-Based Algorithm for Discovering Clusters in Large Spatial Databases with NoiseDBSCAN
  9. Visualizing Data using t-SNEt-SNE 降维
  10. UMAP: Uniform Manifold Approximation and ProjectionUMAP 降维
  11. A Global Geometric Framework for Nonlinear Dimensionality ReductionIsomap
  12. Nonlinear Dimensionality Reduction by Locally Linear EmbeddingLLE
  13. On Spectral Clustering: Analysis and an Algorithm谱聚类
  14. Latent Dirichlet AllocationLDA 主题模型
  15. Learning the Parts of Objects by Non-negative Matrix Factorization非负矩阵分解
  16. Fast Algorithms for Mining Association RulesApriori 关联规则
  17. The PageRank Citation Ranking: Bringing Order to the WebPageRank
  18. Regression Shrinkage and Selection via the LassoLASSO
  19. Regularization and Variable Selection via the Elastic NetElastic Net
  20. Reducing the Dimensionality of Data with Neural Networks自编码器
  21. Dropout: A Simple Way to Prevent Neural Networks from OverfittingDropout
  22. Batch Normalization: Accelerating Deep Network Training批归一化
  23. Adam: A Method for Stochastic OptimizationAdam 优化器
  24. Gradient-Based Learning Applied to Document RecognitionLeNet 卷积网络
  25. ImageNet Classification with Deep Convolutional Neural NetworksAlexNet
  26. Very Deep Convolutional Networks for Large-Scale Image RecognitionVGG
  27. Going Deeper with ConvolutionsInception / GoogLeNet
  28. Deep Residual Learning for Image RecognitionResNet
  29. Densely Connected Convolutional NetworksDenseNet
  30. U-Net: Convolutional Networks for Biomedical Image SegmentationU-Net
  31. Fully Convolutional Networks for Semantic SegmentationFCN
  32. Mask R-CNN实例分割
  33. You Only Look Once: Unified, Real-Time Object DetectionYOLO
  34. Faster R-CNN: Towards Real-Time Object Detection两阶段目标检测
  35. Generative Adversarial NetsGAN
  36. Unsupervised Representation Learning with Deep Convolutional Generative Adversarial NetworksDCGAN
  37. Wasserstein GANWGAN
  38. Auto-Encoding Variational BayesVAE
  39. Sequence to Sequence Learning with Neural NetworksSeq2Seq
  40. Neural Machine Translation by Jointly Learning to Align and Translate注意力机制
  41. Attention Is All You NeedTransformer
  42. BERT: Pre-training of Deep Bidirectional TransformersBERT
  43. Language Models are Unsupervised Multitask LearnersGPT-2
  44. Human-level Control through Deep Reinforcement LearningDQN
  45. Asynchronous Methods for Deep Reinforcement LearningA3C
  46. Proximal Policy Optimization AlgorithmsPPO
  47. Mastering the Game of Go with Deep Neural Networks and Tree SearchAlphaGo
  48. Efficient Estimation of Word Representations in Vector Spaceword2vec
  49. GloVe: Global Vectors for Word RepresentationGloVe
  50. Long Short-Term MemoryLSTM
  51. Learning Phrase Representations using RNN Encoder-DecoderGRU
  52. EfficientNet: Rethinking Model Scaling for Convolutional Neural NetworksEfficientNet
  53. An Image is Worth 16x16 WordsVision Transformer
  54. Learning Transferable Visual Models From Natural Language SupervisionCLIP
  55. Denoising Diffusion Probabilistic ModelsDDPM 扩散模型

C++ 开发资料

聚焦语言能力、现代工程工具链、测试与性能诊断。

  1. C++ Core Guidelines现代 C++ 设计与安全实践
  2. C++ Working Draft标准草案阅读入口
  3. cppreference: LanguageC++ 语言特性索引
  4. CMake Tutorial跨平台构建基础
  5. C++ Modules模块化编译
  6. C++ Ranges范围库与组合式算法
  7. RAII资源生命周期管理
  8. C++ Memory Library智能指针与内存工具
  9. C++ Coroutines协程语言支持
  10. Concepts and Constraints泛型约束
  11. AddressSanitizer内存错误诊断
  12. GoogleTest Primer单元测试
  13. Google Benchmark User Guide微基准测试
  14. Boost.Asio网络与异步 I/O
  15. fmt Documentation现代格式化库

Rust 开发资料

以官方语言资料为主,覆盖异步、工程化、服务端与并发实践。

  1. The Rust Programming Language官方入门书
  2. The Rust Reference语言语义参考
  3. The RustonomiconUnsafe Rust 深入资料
  4. Asynchronous Programming in Rust异步编程
  5. The Cargo Book构建、依赖与发布
  6. Clippy Documentation静态检查与代码改进
  7. Rust API Guidelines库 API 设计
  8. Tokio Tutorial异步运行时
  9. axum DocumentationWeb 服务框架
  10. Serde序列化与反序列化
  11. Rust By Example示例驱动学习
  12. The Rust Performance Book性能分析与优化
  13. The Embedded Rust Book嵌入式开发
  14. Rust Design Patterns常见模式与反模式
  15. Rust Atomics and Locks并发原语与同步

数据库与分布式系统资料

从关系模型和存储引擎,到全球分布式事务与现代分析系统。

  1. A Relational Model of Data for Large Shared Data Banks关系模型
  2. Organization and Maintenance of Large Ordered IndexesB-tree
  3. System R: Relational Approach to Database Management查询优化与关系数据库
  4. ARIES: A Transaction Recovery Method事务恢复
  5. The Log-Structured Merge-TreeLSM-tree
  6. Bigtable: A Distributed Storage System for Structured DataGoogle 分布式存储论文
  7. MapReduce: Simplified Data Processing on Large ClustersGoogle 批处理论文
  8. Dynamo: Amazon's Highly Available Key-value Store高可用 KV 存储
  9. Cassandra: A Decentralized Structured Storage System去中心化宽表
  10. Spanner: Becoming a Global Database全球一致性事务
  11. F1: A Distributed SQL Database That ScalesGoogle 分布式 SQL
  12. In Search of an Understandable Consensus AlgorithmRaft 共识
  13. The Part-Time ParliamentPaxos 共识
  14. ClickHouse: A High-Performance Column-Oriented DBMS列式分析数据库
  15. Apache Iceberg Specification数据湖表格式

算力与硬件观察

面向本地开发、AI 实验和数据服务的周期性技术栏目。型号、驱动和价格变化很快:购买前应重新核对官方规格、平台兼容性、功耗与实际工作负载。

GPU / NVIDIA RTX

RTX 产品线怎么区分

  • GeForce RTX 50 Series面向个人工作站、创作与本地 AI 实验;先比较显存容量、散热、供电和机箱空间。
  • GeForce RTX 官方规格对比以同一官方页面比较显存、代际和核心规格,避免只看单一跑分。
  • NVIDIA RTX PRO专业图形与 AI 工作站路线;适合对专业驱动、显存或认证软件有明确要求的场景。

GPU COMPUTING / CUDA

CUDA 技术入口

  • CUDA Toolkit编译器、运行时、库、调试与性能分析工具的官方下载与资料入口。
  • CUDA Programming Guide先理解线程层级、内存层级和异步执行模型,再写 kernel。
  • CUDA Best Practices Guide以 profiler 测量为准,循序检查访存、并行度和数据传输。

MEMORY / STORAGE

内存与存储的判断顺序

建议的未来购买时间线

现在 — 2026 Q4

先完成基线

记录 CPU、内存、GPU、磁盘占用、I/O 等待和训练/推理耗时;只有瓶颈能被数据复现,升级才有可验证的收益。

2027 H1

按 AI 任务评估首张 GPU

当本地模型实验已形成稳定任务、CPU 或云端排队成为主要阻碍时,按模型所需显存选择单卡 RTX;下单前核对 PCIe 槽位、供电、散热与驱动兼容性。

2027 H2

扩展热数据与内存

当热数据盘长期超过约 70% 可用容量、I/O 等待持续影响任务,或工作集频繁触发 swap/外存时,再以容量和耐久度优先扩展 NVMe 或 DDR5。

每年一次

复核冷存储与恢复

按数据增长、保留期和恢复演练结果规划 HDD/备份扩容;先验证恢复,再讨论增加原始容量。

以上是复核窗口和触发条件,不是固定采购清单或价格建议。产品可得性、保修、价格、驱动版本和平台兼容性应在下单当周重新确认。