Skip to main content
ScrayleScrayle
scheduled scrapingself-healing selectorscron web scrapingworkflow schedulingdata delivery webhooks

Scheduled Workflows, Self-Healing Selectors & Data Delivery

Schedule scraping jobs with cron expressions, let AI fix broken selectors automatically, and push results to webhooks or email - no polling required.

Scrayle Team25-Jul-20259 min read

Extracting data once is the easy part. The real value of web scraping comes from monitoring sources continuously - tracking price changes, watching competitor inventory, aggregating news feeds, or collecting leads on a regular cadence. But running scrapers reliably on a schedule introduces a set of problems that most teams underestimate: cron job management, stale selectors, error recovery, data delivery, and monitoring.

Scrayle's workflow engine handles all of this. You define a scraping workflow (either by writing it manually, using the visual builder, or letting the AI agent generate one), attach a schedule, and configure where results should go. Scrayle runs it on time, heals broken selectors automatically, retries on transient failures, and delivers the data where you need it.

Anatomy of a Workflow

A Scrayle workflow is a JSON document that describes a sequence of browser actions and extraction rules. Here's a simplified example:

{
"name": "HN Top Stories",
"steps": [
{
"action": "navigate",
"url": "https://news.ycombinator.com"
},
{
"action": "wait",
"selector": ".athing",
"timeout": 10000
},
{
"action": "extract",
"items": {
"selector": ".athing",
"limit": 30,
"fields": {
"title": ".titleline > a",
"url": { "selector": ".titleline > a", "attribute": "href" },
"rank": ".rank"
}
}
}
]
}

Workflows support a rich set of actions: navigate, click,type, scroll, wait, extract,screenshot, evaluate (run arbitrary JavaScript), andcondition (branching logic). You can loop over pages, handle pop-ups, and chain multiple extraction steps.

Scheduling with Cron Expressions

Cron schedules are configured in the Scrayle dashboard under Workflows → Schedule. Select a workflow, choose a cron expression and timezone, and Scrayle runs it automatically. The SDK lets you trigger on-demand runs and check results programmatically:

import { Scrayle } from '@scrayle/sdk';
const client = new Scrayle({ apiKey: process.env.SCRAYLE_API_KEY });
// Trigger an immediate run of an existing workflow
const run = await client.workflows.run({
workflowId: 'wf_x8k2m9',
workflowName: 'HN Top Stories',
steps: [], // Step definitions are stored with the workflow
selfHeal: true,
browserMode: 'stealth',
});
// Wait for completion and inspect the result
const result = await client.workflows.waitForCompletion(run.taskId);
console.log('Status:', result.status);

Common cron patterns:

  • 0 9 * * 1-5 - Weekdays at 9 AM
  • 0 */6 * * * - Every 6 hours
  • 0 0 * * 0 - Weekly on Sunday at midnight
  • */30 * * * * - Every 30 minutes
  • 0 9 1 * * - First of every month at 9 AM

Scrayle evaluates cron expressions in the timezone you specify (defaults to UTC). Scheduled runs are queued and executed with a guaranteed start window of +/- 60 seconds from the scheduled time.

Self-Healing Selectors

This is the feature that saves the most engineering time in production. Here's how it works:

  1. A scheduled workflow runs and attempts to execute its selectors against the current page
  2. If a selector fails (returns no elements when it previously returned results), the workflow engine flags it as a selector failure
  3. The engine takes a screenshot of the current page state and sends it to the AI, along with the original selector, the expected data shape, and the page's DOM
  4. The AI analyzes the page and generates a new selector that targets the same data
  5. The workflow retries with the new selector. If it succeeds, the selector is permanently updated in the workflow definition
  6. You receive a notification showing the old and new selectors so you can review what changed

Self-healing covers the most common scraping breakage scenarios:

  • Class name changes: When .product-card-v2 becomes .product-card-v3 or a hashed class like .css-1a2b3c
  • DOM restructuring: When the HTML nesting changes but the visual layout stays similar
  • Attribute changes: When data-testid values are updated
  • Partial changes: When some selectors still work but others are broken - only the broken ones are healed

In our internal testing, self-healing successfully resolved approximately 85% of selector breakages without human intervention. The remaining 15% were major site redesigns that changed the fundamental data structure.

Data Delivery

Getting data out of Scrayle and into your systems is just as important as extracting it. Workflows support multiple delivery targets that run automatically after each successful execution:

Webhook

POST the results as JSON to any HTTP endpoint. Scrayle includes retry logic with exponential backoff for failed deliveries. Configure a webhook delivery target in the dashboard under Workflows → Schedule → Delivery:

{
"webhook": {
"url": "https://your-api.com/ingest",
"headers": { "Authorization": "Bearer your-token" },
"template": {
"source": "scrayle",
"timestamp": "{{run.timestamp}}",
"items": "{{data}}"
}
}
}

Email

Receive results as an attached CSV or JSON file. Useful for non-technical stakeholders who need regular data reports.

delivery: {
email: {
subject: 'Weekly Price Report - {{run.date}}',
format: 'csv', // or 'json'
},
}

Scrayle Storage

Store results directly in Scrayle's built-in storage system. Each workflow run creates a new entry in a dataset, making it easy to query historical data and track changes over time.

delivery: {
storage: {
dataset: 'price-tracking',
// Results are automatically appended with run metadata
},
}

S3-Compatible Storage

Push files directly to Amazon S3, Google Cloud Storage, Cloudflare R2, or any S3-compatible object store.

delivery: {
s3: {
bucket: 'my-scraping-data',
region: 'us-east-1',
prefix: 'scrayle/prices/',
// Files are named: {prefix}{workflow-id}/{run-date}.json
credentials: {
accessKeyId: process.env.AWS_ACCESS_KEY_ID,
secretAccessKey: process.env.AWS_SECRET_ACCESS_KEY,
},
},
}

You can configure multiple delivery targets for the same workflow. For example, send data to a webhook for real-time processing while also storing a copy in Scrayle storage for historical analysis.

Error Handling and Retries

Production scrapers fail. Pages go down, CAPTCHAs can't be solved, network connections drop, and content takes too long to load. Scrayle's workflow engine handles this with a configurable retry policy:

{
"retry": {
"maxAttempts": 3,
"backoffMs": 30000,
"backoffMultiplier": 2
},
"onFailure": {
"webhook": "https://your-api.com/alerts",
"email": ["[email protected]"]
}
}

Retry and failure notification settings are configured in the dashboard under Workflows → Schedule → Error Handling.

Each retry gets a fresh browser session with a new IP address and fingerprint, which handles the common case where a failure was caused by the specific IP being blocked.

Monitoring and Run History

Every workflow run is logged with detailed telemetry:

  • Run status: success, failed, partially succeeded (some data extracted but not all)
  • Duration: Total browser time and per-step timing
  • Data volume: Number of items extracted, byte size
  • Screenshots: Captured at key points during execution
  • Selector health: Which selectors succeeded, which were healed, which failed
  • Delivery status: Whether each delivery target received the data

You can view run history in the dashboard or query it programmatically:

// List the 10 most recent runs across all workflows
const runs = await client.workflows.listRuns(10);
for (const run of runs) {
console.log(run.taskId, run.status, run.startedAt);
}
// Check status of a specific run by task ID
const status = await client.workflows.getRunStatus('task_abc123');
console.log(status.status, status.error);

Building Workflows in the Visual Editor

If you prefer a visual approach, the Scrayle dashboard includes a drag-and-drop workflow editor. You can:

  • Add steps by dragging action blocks onto a canvas
  • Configure each step with a form interface (no JSON editing)
  • Test individual steps against a live browser preview
  • See selector matches highlighted on the page in real-time
  • Set up scheduling and delivery without writing code

The visual editor works on any workflow - including those generated by the AI agent. A typical pattern is to let the AI generate the initial workflow, then use the visual editor to fine-tune selectors, add error handling, or modify the output format.

Putting It All Together

Here's a complete example that creates a price monitoring pipeline from scratch:

The full pipeline has three parts: build the workflow in the dashboard using the AI agent, configure the schedule and delivery targets in the dashboard, then use the SDK for on-demand runs and status checks.

Step 1 - Build the workflow (dashboard AI agent):

Go to https://example-store.com/laptops
Extract all laptops with:
- product name
- current price
- original price (if on sale)
- availability
Get results from the first 3 pages

Step 2 - Configure schedule & delivery (dashboard):

{
"schedule": { "cron": "0 9,21 * * *", "timezone": "UTC" },
"delivery": {
"storage": { "dataset": "laptop-prices" },
"webhook": { "url": "https://your-api.com/price-alerts" },
"email": { "to": ["[email protected]"], "format": "csv" }
},
"retry": { "maxAttempts": 3, "backoffMs": 60000 }
}

Step 3 - Trigger and monitor via SDK:

import { Scrayle } from '@scrayle/sdk';
const client = new Scrayle({ apiKey: process.env.SCRAYLE_API_KEY });
// Trigger an on-demand run (e.g. for testing)
const run = await client.workflows.run({
workflowId: 'wf_laptop_prices',
workflowName: 'Laptop Price Monitor',
steps: [],
selfHeal: true,
browserMode: 'stealth',
});
console.log('Queued task:', run.taskId);
// Wait for the result
const result = await client.workflows.waitForCompletion(run.taskId);
console.log('Status:', result.status);

Once this is running, you have a fully automated pipeline that:

  • Extracts laptop prices twice a day
  • Stores every data point for historical analysis
  • Sends real-time updates to your API for price-drop alerts
  • Emails a weekly CSV summary to your purchasing team
  • Automatically fixes broken selectors when the site updates
  • Alerts your ops team if something goes permanently wrong

Best Practices for Production Workflows

  • Start with the AI agent, then optimize. Let AI generate the first version of your workflow. It handles 90% of the selector work. Then manually refine any selectors that are overly broad or fragile.
  • Use descriptive workflow names. You'll accumulate many workflows over time. Names like "laptop-prices-amazon-us" are much better than "my-scraper-2".
  • Set reasonable schedules. Don't scrape more frequently than you need the data. Hourly updates on a page that changes weekly wastes credits and puts unnecessary load on the target site.
  • Monitor selector health. Even with self-healing, review healed selectors periodically. Multiple heals in short succession may indicate the site is actively restructuring.
  • Use Scrayle storage for data you want to query later. Webhooks are great for real-time pipelines, but having a stored copy lets you backfill, debug, and analyze without re-running the workflow.
  • Test workflows manually before scheduling. Run a workflow once via the SDK or dashboard to verify it extracts the right data before committing to a schedule.

What's Next


Questions? Check the documentation for detailed API references, or open a support ticket in the dashboard.