Tag: Finance

  • Beyond the Bell Curve: Mutual information based stock networks

    Beyond the Bell Curve: Mutual information based stock networks

    In traditional quantitative finance, the structural failure of the Pearson correlation under non-Gaussian conditions often remains a massive blind spot. Most investors treat the market as a static spreadsheet of linear relationships, assuming that if Stock A moves, Stock B follows in a predictable straight line. However, when we zoom into the temporal dynamics of information flow—specifically at the 30-second “tick” level—the market reveals itself as a living, breathing network. By applying Random Matrix Theory (RMT) as a baseline for “noise dressing” and moving toward “Mutual Information” (MI), we can see the market’s hidden architecture. Derived from information theory, MI allows us to detect the “real” complex dependencies—both linear and non-linear—that traditional 20th-century models ignore. In emerging markets like India, where interactions are structurally stronger than in developed markets, this high-definition map is not just a theoretical luxury; it is a necessity for survival. In this article I am summarizing key findings from my research article jointly written by Prof Amber Habib, Professor Shiv Nadar University (https://doi.org/10.1371/journal.pone.0221910)

    The Structural Failure of the Pearson Correlation

    The most common tool in the strategist’s kit, the correlation coefficient, is fundamentally ill-equipped for high-frequency environments. Its primary flaw is its inability to detect non-linear dependencies. When we analyze the spectrum of the correlation matrix, we often find that while high correlation aligns with high MI, there is a substantial “non-linear data”—instances where stocks exhibit low correlation but high mutual information. As the source research emphasizes:  “If there is a non-linear relationship, the correlation coefficient may fail to capture it.”  For a quant, relying on linear-only thinking creates a dangerous false sense of security. If your risk models don’t account for these invisible threads, you are effectively flying blind through market turbulence, missing the deep dependencies that emerge when Gaussian assumptions break down.

    Speed Changes the Physics of Finance

    The rules of information flow shift dramatically based on the time scale. When analyzing daily returns, the discrepancy between linear correlation and mutual information is relatively stable. However, the “physics” changes at the high-frequency level. In 30-second intervals, non-linear interactions become significantly more pronounced. This suggests that market “noise” at the intraday level isn’t just random; it’s a dense web of rapid, non-linear bursts. For a strategist, this means a model optimized for long-term daily returns might be disastrous if applied to an algorithmic intraday strategy. At high frequencies, the market is less about fundamental valuation and more about the mechanics of how information propagates through the network.

    Systemic Fragility and the “Election Effect”

    We worked with 2014 data and observed that a major political events, such as the 2014 Indian general election, do more than just increase volatility—they reshape the entire topology of the stock network. During periods of political uncertainty, we observe a phenomenon where “common expectations” cause the market to move with collective intensity. From a network perspective, the scale-free property of the market undergoes a radical shift. In “normal” times, the power law exponent (α \alpha ) typically sits between 2 and 3. During the election, however,  α\alpha  dropped below 2. This drop indicates the emergence of a “thicker tail” and a proliferation of high-degree “hubs”—specifically in the  Financial Services and IT sectors .“During election time investors develop common expectations… market becomes volatile and at the same time stronger and more number of interactions are observed amongst the stocks. “When  α<2\alpha < 2“, systemic fragility spikes. The system-wide spike in mutual information signals a collapse of diversification benefits, as the “common expectation” forces almost every stock to react to the same central information hubs.

     Optimizing for the “Periphery”

    To visualize the market’s core, we use a Minimum Spanning Tree (MST)—a “skeleton” of the network that retains only the strongest connections while stripping away redundant loops. Within this skeleton, some stocks act as “hubs” (like ICICI Bank or PNB), while others sit on the “periphery”. Counter-intuitively, the boring periphery is where the intraday profit hides. Hubs possess high Eigenvector Centrality, meaning they are connected to other highly connected nodes. While this makes them central to information flow, it also makes them informationally overloaded and highly susceptible to systemic contagion. Conversely, peripheral stocks have low Eigenvector Centrality; their neighbors are also not central. By staying away from the center of the market noise, traders can find “pure” signals and better risk-adjusted returns, as these peripheral assets are more isolated from the chaotic fluctuations of the primary hubs.

    Precision over Breadth in Portfolio Selection

    A hallmark of modern portfolio theory is that broader is better. However, the data reveals a more nuanced reality for high-frequency environments. While a full 45-stock Markowitz model (maximizing the Sharpe ratio) performs best in absolute terms of the  return-to-stability ratio, it is often impractical for the “human” intraday trader or focused algorithms. The real breakthrough lies in small-cap constraints. When selecting small portfolios of 3, 5, or 10 stocks, MI-based peripheral selection significantly outperforms traditional correlation-based methods. Using entropy-based measures to pick a handful of peripheral stocks provides a more stable, efficient outcome than trying to manage the “clutter” of the entire CNX100. You don’t need to track 89 stocks to win; you just need to identify the 5 that are most informationally distinct from the noisy center.

    Conclusion: The Future is Entropic

    The 20th-century reliance on linear models is giving way to a 21st-century paradigm rooted in Information Theory. This case study of the Indian National Stock Exchange (NSE) proves that non-linearities are a hallmark of high-frequency trading, especially in emerging markets where interactions are more intense. As we look toward the future, the question for every quantitative strategist is simple: are you measuring the “real” complex connections in your portfolio using tools like Mutual Information and Adjacency Matrices, or are you still relying on the visible, linear illusions of the bell curve?

  • Beyond Correlation: How AI and Physics-Based Synchronization Are Decoding the Indian Stock Market’s Hidden Chaos

    Beyond Correlation: How AI and Physics-Based Synchronization Are Decoding the Indian Stock Market’s Hidden Chaos

    Traditional financial tools often struggle to capture the true complexity of how stocks move together. While many investors rely on simple linear correlation to understand market relationships, these standard methods frequently fail to account for the “messy,” non-linear, and lagged co-movements that define real-world trading. When markets turn volatile, these surface-level metrics often break down, leaving portfolios exposed to risks that traditional models never saw coming. One of my recent research with my student  Mr. Sanjay Sathish from the Shiv Nadar Institution of Eminence also presented at the 34th European Conference on Operational Research, offers a more sophisticated approach. By analyzing 21 years of Indian stock market data (2002–2023), we tried to move beyond simple price trends to train Artificial Intelligence on the concept of “synchronization.” This method treats the market as a dynamic system, utilizing recurrence-based techniques to find patterns hidden in the noise of over two decades of trading history.

    Beyond Correlation—The Non-Linear Reality of the Market

    Traditional finance typically views stock relationships through a linear lens, assuming that if Stock A moves, Stock B will move proportionally. However, the researchers argue that these methods are insufficient because they miss the nuances of non-linear or lagged movements where one stock might react to another after a significant delay or in a complex, non-proportional manner. This “non-linear reality” means that stocks can be deeply connected even when their correlation coefficients appear low. Understanding these hidden synchronization patterns is a game-changer for modern portfolio construction. By identifying how stocks truly move in tandem across different market cycles, investors can develop more robust predictive models that account for systemic “synchronicity” rather than just surface-level trends. As noted in the research motivation: “Understanding of the price movements can help in building better predictive models which can be used to construct profitable portfolios.”

    Visualizing Chaos with Multidimensional Cross Recurrence Plots (CRPs)

    To capture these complex dynamics, the study employed Cross Recurrence Plots (CRPs), a technique originating in the study of dynamical systems in physics. Unlike standard charts, the researchers used a multidimensional approach, analyzing a four-part time series for each pair: the daily closing Price and Volume for Stock A, and the daily closing Price and Volume for Stock B. To ensure the data was embedded correctly into a higher-dimensional space, we utilized the “False Nearest Neighbors” method to determine the optimal consistency across the 20-stock universe.The study focused on 20 highly-capitalized stocks across 14 different sectors, generating 190 unique pairs for analysis. This method is particularly powerful because it can detect “phase transitions and critical points”—sudden, sharp changes in market behavior that standard linear charts often ignore. By calculating the actual distances between points in a phase diagram, the CRPs extract the “synchronization dynamics” that define how different sectors interact under pressure.

    Why RNNs and LSTMs are the “Memory” the Market Needs

    While CRPs provide the mathematical framework, the research utilized Deep Learning to predict future synchronization states. Recurrent Neural Networks (RNNs) and Long Short-Term Memory (LSTM) models were chosen for their ability to “remember” patterns in sequential data. Before feeding the data into the AI, we carried out a critical preprocessing step: we averaged the distance matrices into “mean of 3×3 blocks” to make the high-dimensional data manageable for the neural networks. The LSTM model architecture was meticulously designed for precision and stability, featuring:

    • Two LSTM layers:  These utilize the  tanhtanh activation function  to process the extracted distance sequences.
    • Dropout layers:  Set at a rate of 0.2, these provide regularization to prevent the model from over-fitting to historical noise.
    • Adam optimizer:  Used to efficiently minimize error during the training of over 3,600 data points. To evaluate success, metrics such as MAE, RMSE, R-square, and MAPE were used. The precision of the model is most evident in the “actual vs. predicted” distance graphs, where the predicted yellow lines track the actual green lines with remarkable fidelity. Most impressively, the LSTM was able to accurately mirror the  chaotic spikes  in distance—those specific moments where two stocks suddenly diverged or “desynchronized” before returning to their normal state.
    image

    Takeaway 4: The Binary “Sync” Signal—Classifying Market States

    The final step of the research transforms complex distance data into actionable signals for investors. By using the Heaviside function, we converted continuous distance measures into binary labels:  Synchronous (1)  or  Non-synchronous (0) . A distance below a specific threshold indicates the stocks are moving in lockstep, while a distance above it suggests they are acting independently. We tested various thresholds, known as  Recurrence Rates (RR) , ranging from 20% to 45%. A lower recurrence rate represents a more stringent definition of “sync,” while a higher rate is more inclusive of looser movements. To determine which threshold was most reliable, we used the  Coefficient of Variance (CV)  to rank the models; lower CV values indicated the most consistent performance across the 21-year dataset. This classification turns complex market noise into a clear signal, allowing traders to see exactly when the fundamental synchronization of a pair has broken down.