We reproduce chain-of-thought experiments across base, instruction-tuned, and reasoning models. The results suggest that reported CoT gains on modern models are mostly artifacts of suppression, not necessarily reasoning improvements.
In this post, we discuss the issue of checkerboard artefacts introduced by transpose convolution layers as well as proposed solutions to this problem. Experiments on audio synthesis are also given, further motivating these methods.
An exploration of Noise Contrastive Estimation (NCE) and how it enables Contrastive Predictive Coding (CPC) for learning useful representations from high-dimensional sequential data like speech.