new

Get trending papers in your email inbox!

Subscribe

Daily Papers

byAK and the research community

Oct 8

Extending the Design Space of Graph Neural Networks by Rethinking Folklore Weisfeiler-Lehman

Message passing neural networks (MPNNs) have emerged as the most popular framework of graph neural networks (GNNs) in recent years. However, their expressive power is limited by the 1-dimensional Weisfeiler-Lehman (1-WL) test. Some works are inspired by k-WL/FWL (Folklore WL) and design the corresponding neural versions. Despite the high expressive power, there are serious limitations in this line of research. In particular, (1) k-WL/FWL requires at least O(n^k) space complexity, which is impractical for large graphs even when k=3; (2) The design space of k-WL/FWL is rigid, with the only adjustable hyper-parameter being k. To tackle the first limitation, we propose an extension, (k,t)-FWL. We theoretically prove that even if we fix the space complexity to O(n^k) (for any kgeq 2) in (k,t)-FWL, we can construct an expressiveness hierarchy up to solving the graph isomorphism problem. To tackle the second problem, we propose k-FWL+, which considers any equivariant set as neighbors instead of all nodes, thereby greatly expanding the design space of k-FWL. Combining these two modifications results in a flexible and powerful framework (k,t)-FWL+. We demonstrate (k,t)-FWL+ can implement most existing models with matching expressiveness. We then introduce an instance of (k,t)-FWL+ called Neighborhood^2-FWL (N^2-FWL), which is practically and theoretically sound. We prove that N^2-FWL is no less powerful than 3-WL, and can encode many substructures while only requiring O(n^2) space. Finally, we design its neural version named N^2-GNN and evaluate its performance on various tasks. N^2-GNN achieves record-breaking results on ZINC-Subset (0.059), outperforming previous SOTA results by 10.6%. Moreover, N^2-GNN achieves new SOTA results on the BREC dataset (71.8%) among all existing high-expressive GNN methods.

  • 7 authors
·
Jan 13, 2024

DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence

We present a preview version of DeepSeek-V4 series, including two strong Mixture-of-Experts (MoE) language models -- DeepSeek-V4-Pro with 1.6T parameters (49B activated) and DeepSeek-V4-Flash with 284B parameters (13B activated) -- both supporting a context length of one million tokens. DeepSeek-V4 series incorporate several key upgrades in architecture and optimization: (1) a hybrid attention architecture that combines Compressed Sparse Attention (CSA) and Heavily Compressed Attention (HCA) to improve long-context efficiency; (2) Manifold-Constrained Hyper-Connections (mHC) that enhance conventional residual connections; (3) and the Muon optimizer for faster convergence and greater training stability. We pre-train both models on more than 32T diverse and high-quality tokens, followed by a comprehensive post-training pipeline that unlocks and further enhances their capabilities. DeepSeek-V4-Pro-Max, the maximum reasoning effort mode of DeepSeek-V4-Pro, redefines the state-of-the-art for open models, outperforming its predecessors in core tasks. Meanwhile, DeepSeek-V4 series are highly efficient in long-context scenarios. In the one-million-token context setting, DeepSeek-V4-Pro requires only 27% of single-token inference FLOPs and 10% of KV cache compared with DeepSeek-V3.2. This enables us to routinely support one-million-token contexts, thereby making long-horizon tasks and further test-time scaling more feasible. The model checkpoints are available at https://huggingface.co/collections/deepseek-ai/deepseek-v4.

deepseek-ai DeepSeek
·
Apr 25 5