Large Language Models (LLMs) have evolved beyond being mere downstream Natural Language Processing (NLP) components: an increasing number of recent LLM-based systems can perform end-to-end planning, assembly, and evaluation of machine learning (ML) workflows. Advances in alignment techniques, tool integration, and program-generation methods enable LLMs to create end-to-end automated pipelines that execute from data collection through model deployment, verification, and reporting. The review examines three essential areas of current research, which include (i) the evolution of Transformer-based LLMs with tool-use interfaces, (ii) the exploration of fundamental studies about LLM capabilities, risks, and evaluation methodologies, and (iii) the design of new systems for LLM-based Automated Machine Learning (AutoML) and pipeline management systems that use multi-agent architecture, compiler frameworks, and iterative improvement cycles. The evaluation compares different approaches through their planning methods, verification techniques, optimization systems, monitoring capabilities, and operational expenses. The case studies report significant speed improvements, but researchers must address four essential problems that affect the reliable deployment of systems. The review follows IEEE Access guidelines for articles by including a clear methods section that explains the sources, time frame, and selection criteria, and by presenting visual representations that link systems to their corresponding pipeline stages. The review presents a unified, pipeline-stage–aligned conceptual framework that integrates terminology, design rules, and failure-detection practices for analyzing and validating LLM-based end-to-end ML systems. As a narrative review, this work emphasizes system-level design patterns and representative benchmarks rather than exhaustive quantitative meta-analysis.