Histopathological examination remains the gold standard for oral squamous cell carcinoma (OSCC) diagnosis, yet manual slide review is time-consuming and subject to inter-observer variability. Existing deep learning approaches require exhaustive image-level or pixel-level annotations, which are prohibitively expensive at scale. This study introduces OralHistoMIL, a weakly-supervised Attention-Based Multi-Instance Learning (AB-MIL) framework that identifies diagnostically significant regions using only slide-level labels, eliminating the need for detailed region annotations. Each image is decomposed into overlapping 224x224 patches via sliding window extraction with tissue content filtering. A frozen ResNet-50 backbone extracts 2048-dimensional features per patch, projected to a 512-dimensional embedding space. A gated attention mechanism jointly learns tanh and sigmoid pathways, producing instance-level importance scores that identify which patches contribute most to the diagnosis. The attention-weighted bag representation is classified as Normal versus OSCC. We evaluated OralHistoMIL on 1,224 histopathological images (290 normal, 934 OSCC) from the Mendeley Oral Cancer Imaging Database across two magnifications (100x and 400x). Despite requiring only slide-level supervision, AB-MIL achieved accuracy of 0.922 ± 0.012, F1-score of 0.924 ± 0.018, and AUC-ROC of 0.970 ± 0.015, with sensitivity of 0.930 and specificity of 0.897, performing comparably to fully-supervised baselines including DenseNet-121 (AUC=0.978) and EfficientNet-B0 (AUC=0.978) while providing spatial interpretability unavailable to standard classifiers. Attention heatmap analysis confirmed that high-attention patches corresponded to nuclear pleomorphism, stromal invasion, and altered epithelial architecture, aligning with established OSCC diagnostic criteria and offering pathologists transparent visual explanations without exhaustive annotation.