You built a classification model to predict whether the S&P 500 will go up or down tomorrow. Standard 5-fold cross validation shows 65% accuracy, but it consistently loses money in live trading. What is fundamentally wrong with your validation approach?