Automated offensive language detection (OLD) is essential for online safety, yet remains particularly challenging for under-resourced languages and their dialectal variations. This paper addresses these challenges by evaluating OLD across three low-resource Arabic dialects: Egyptian, Libyan, and Levantine, using newly collected dialect-specific datasets and a novel parallel corpus for controlled cross-dialect analysis.We benchmark a wide range of models, including encoder-only, encoder–decoder, and decoder-only Large Language Models (LLMs), alongside traditional machine learning baselines (XGBoost, LightGBM), under zero-shot, few-shot, and fine-tuned settings. Our experiments reveal that task-specific fine-tuning consistently delivers the highest performance, especially for models pre-trained on Arabic corpora such as AraBERT, while few-shot learning yields modest and sometimes unstable gains. Surprisingly, classical models leveraging AraBERT embeddings remain competitive in zero-shot settings. Cross-dialect evaluations further highlight substantial performance variability driven by dialectal distance, lexical ambiguity, and contextual offensiveness. These findings offer a comprehensive benchmark for OLD in dialectal Arabic and provide practical guidance for developing robust moderation systems in linguistically diverse, low-resource environments.