I am a Data Engineer with 3+ years of experience specializing in high-performance data storage and automated pipeline orchestration. I focus on optimizing ETL/ELT workflows and leveraging OLAP databases to turn large-scale logs into actionable insights.
| Category | Tools & Technologies |
|---|---|
| Languages | Python, Java, SQL, Shell Script |
| Data Engineering | Apache Airflow, Spark, Hadoop Ecosystem |
| Databases | Clickhouse, MongoDB, Redis, MySQL, MariaDB |
| Development | Docker, GitLab CI/CD, Linux (Ubuntu) |
-
High-Performance Data Warehousing
-
Optimized Clickhouse storage solutions, implementing effective partitioning and indexing strategies to handle terabytes of historical logs.
-
Significantly reduced query latency for high-frequency data retrieval, enabling real-time analytical capabilities.
-
Scalable Data Pipeline Orchestration
-
Architected and managed complex DAGs via Apache Airflow, automating the full data lifecycle from raw ingestion to production-ready reporting layers.
-
Integrated Spark for distributed processing, achieving a 40%+ improvement in transformation speed compared to legacy systems.
-
Backend Automation & System Integration
-
Developed robust backend services and API integrations to bridge diverse data sources and improve system interoperability.
-
Leveraged Redis for high-speed caching and Docker for consistent service containerization across environments.
-
Development Workflow Optimization
-
Established GitLab CI/CD pipelines to automate testing and deployment, ensuring high reliability of data services and minimizing production risks.
- LinkedIn: linkedin.com/in/marcus-lin
- Email: s09203647@gmail.com
- Location: Taipei, Taiwan 🇹🇼
I am focused on Database Internals and Distributed Systems. Currently exploring advanced data modeling techniques and enhancing system observability for large-scale data infrastructures.
