SFTP to Azure Data Lake Storage Gen2 Data Migration Using Azure Data Factory

SFTP to Azure Data Lake Storage Gen2 Data Migration Using Azure Data Factory

Introduction

Migrating files from an SFTP server to Azure Data Lake Storage Gen2 (ADLS Gen2) is a common cloud data-integration requirement. Azure Data Factory (ADF) provides a managed and scalable way to securely transfer files without developing a custom application.

The basic architecture is:

1. Prerequisites

Before starting, ensure you have:

  • Azure subscription
  • SFTP server hostname and port
  • SFTP username and authentication credentials
  • Azure Data Factory
  • ADLS Gen2 storage account
  • Required Azure RBAC permissions
  • Network connectivity between ADF and the SFTP server 

For production, use SSH key authentication where supported and avoid storing credentials directly in pipelines.

2. Create ADLS Gen2 Storage

Create an Azure Storage Account and enable:

Hierarchical namespace = Enabled

This enables ADLS Gen2 capabilities.

Create a container, for example:

datalake

Recommended folder structure:

datalake/
│
├── raw/
│   └── sftp/
│       ├── customers/
│       ├── orders/
│       └── invoices/
│
├── processed/
│
└── archive/

The raw layer should preserve the original source files.

3. Create Azure Data Factory

Create an Azure Data Factory instance and open ADF Studio.

ADF will be responsible for:

  • Connecting to SFTP
  • Reading files
  • Copying files to ADLS
  • Scheduling migration
  • Monitoring execution
  • Handling failures

4. Create SFTP Linked Service

In ADF:

Manage → Linked services → New → SFTP

Configure:

  • Host
  • Port: 22
  • Username
  • Authentication

Use password or SSH key authentication according to your organization's security requirements.

Test the connection before continuing.

Production recommendation

Store sensitive credentials in Azure Key Vault instead of hardcoding them in the pipeline.

5. Create ADLS Gen2 Linked Service

Create another linked service:

Manage → Linked services → New

Select:

Azure Data Lake Storage Gen2

Use Managed Identity where possible and grant ADF the required permissions on the ADLS container/storage account.

6. Create the Copy Pipeline

Create a new ADF pipeline:

PL_SFTP_TO_ADLS

Add a Copy Data activity.

Configure:

Source

SFTP
    ↓
/incoming
For example:
/incoming/customers.csv
/incoming/orders.csv
/incoming/invoices.csv

Destination\

ADLS Gen2
    ↓
/raw/sftp/

For example:

/raw/sftp/customers/customers.csv
/raw/sftp/orders/orders.csv
/raw/sftp/invoices/invoices.csv

7. Incremental File Processing

For recurring migrations, the most important consideration is to avoid processing the same file repeatedly.

A typical flow is:

SFTP
  |
  | New files
  v
ADF
  |
  v
ADLS /raw
  |
  v
Validation
  |
  v
Archive

You can track processed files using:

  • File name
  • File path
  • Last modified date
  • File size
  • Processing status

For larger production implementations, a control table can be used to maintain file-processing history.

8. Scheduling

If files arrive regularly, create an ADF trigger.

Examples:

  • Every 15 minutes
  • Every hour
  • Daily

Choose the schedule based on the business requirement.

For one-time migration, the pipeline can simply be executed manually.

9. Security

Security is mandatory for production implementations.

Recommended approach:

SFTP
   |
   | SSH/SFTP
   v
Azure Data Factory
   |
   +---- Azure Key Vault
   |
   | Managed Identity
   v
ADLS Gen2

Important security considerations:

  • Use secure SFTP connections.
  • Store credentials in Azure Key Vault.
  • Use Managed Identity for Azure resources.
  • Apply least-privilege RBAC.
  • Configure ADLS Gen2 ACLs where required.
  • Restrict network access.
  • Use private networking where required by the organization.

10. Network Connectivity

If the SFTP server is inside an on-premises or private corporate network, ADF may require a Self-hosted Integration Runtime.

Architecture:

Private SFTP Server
       |
       v
Self-hosted Integration Runtime
       |
       v
Azure Data Factory
       |
       v
ADLS Gen2

This is an important consideration before implementing the migration.

11. Validation and Error Handling

A successful pipeline run should not be the only validation.

At minimum, validate:

  • Expected files were transferred.
  • File sizes are correct.
  • Files are readable.
  • No unexpected duplicates were created.

For critical data, checksum or record-count validation can also be implemented.

Configure:

  • Retry
  • Failure handling
  • Logging
  • Alerts

ADF monitoring should be used to identify failed pipelines and activities.

12. Archive Strategy

Do not immediately delete files from the SFTP server after copying them.

A safer approach is:

SFTP /incoming
       |
       v
ADF
       |
       v
ADLS /raw
       |
       v
Validation
       |
       v
SFTP /archive

Only archive or remove the source file after successful transfer and validation, according to the organization's retention policy.

13. Production Architecture

A recommended production architecture is:

14. Mandatory Production Checklist

Before going live, make sure the following are addressed:

  • Connectivity, SFTP connectivity tested
  • Authentication, Secure authentication configured
  • Credentials, Key Vault used for secrets
  • Azure Access, Managed Identity/RBAC configured
  • Storage, ADLS Gen2 with hierarchical namespace
  • Pipeline, Copy Activity tested
  • Incremental Load, Duplicate processing prevented
  • Validation, File transfer validated
  • Error Handling, Retry and failure handling configured
  • Monitoring, ADF monitoring and alerts configured
  • Network, Self-hosted IR/private networking if required
  • Archive, Source-file retention/archiving defined

Conclusion

For most SFTP-to-Azure migrations, Azure Data Factory + ADLS Gen2 provides a simple and scalable architecture:

SFTP
  ↓
Azure Data Factory
  ↓
ADLS Gen2 /raw
  ↓
Validation
  ↓
Archive / Processed
  ↓
Databricks / Synapse / Power BI

The most important production concerns are secure authentication, Key Vault, Managed Identity/RBAC, network connectivity, incremental processing, validation, error handling, monitoring, and source-file retention. Addressing these areas makes the migration reliable, secure, and maintainable.

Post a Comment

0 Comments